Capable of understanding text, audio, vision, video
Industrial-level controllable zero-shot text-to-speech system
Controllable & emotion-expressive zero-shot TTS
Generate audiobooks from e-books, voice cloning & 1107+ languages
Speech recognition module for Python
A nearly-live implementation of OpenAI's Whisper
Official Python inference and LoRA trainer package
Multimodal-Driven Architecture for Customized Video Generation
Instant voice cloning by MIT and MyShell. Audio foundation model
MOSS-TTS-Nano is an open-source multilingual tiny speech generation
Generate audiobooks from e-books
Free, high-quality text-to-speech API endpoint to replace OpenAI
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Fast multimodal LLM for real-time voice interaction and AI apps
Offline Text To Speech synthesis for python
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Interface for OuteTTS models
A lightweight text-to-speech model with zero-shot voice cloning
Label Studio is a multi-type data labeling and annotation tool
Converts text to speech in realtime
A TTS model capable of generating ultra-realistic dialogue
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Use Microsoft Edge's online text-to-speech service from Python
Googles NotebookLM but local
EPUB to audiobook converter, optimized for Audiobookshelf