State-of-the-art TTS model under 25MB
Open-source multi-speaker long-form text-to-speech model
Qwen3-TTS is an open-source series of TTS models
Instant voice cloning by MIT and MyShell. Audio foundation model
MOSS-TTS-Nano is an open-source multilingual tiny speech generation
Converts text to speech in realtime
Towards Human-Sounding Speech
Multi-lingual large voice generation model, providing inference
Open-source framework for intelligent speech interaction
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Long-form streaming TTS system for multi-speaker dialogue generation
PyTorch implementation of VALL-E (Zero-Shot Text-To-Speech)
Conditional Variational Autoencoder with Adversarial Learning