A sound cloning tool with a web interface, using your voice
State-of-the-art TTS model under 25MB
A TTS that fits in your CPU (and pocket)
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Generate audiobooks from e-books
A nearly-live implementation of OpenAI's Whisper
A fast TTS architecture with conditional flow matching
Multi-lingual large voice generation model, providing inference
Miso TTS is an 8 billion, highly emotive text-to-speech model
FAIR Sequence Modeling Toolkit 2
Capable of understanding text, audio, vision, video
VITS2 backbone with multilingual-bert