Capable of understanding text, audio, vision, video
Tokenizer-Free TTS for Multilingual Speech Generation
A nearly-live implementation of OpenAI's Whisper
Qwen3-TTS is an open-source series of TTS models
SOTA Open Source TTS
Controllable & emotion-expressive zero-shot TTS
GLM-4-Voice | End-to-End Chinese-English Conversational Model
TTS model capable of streaming conversational audio in realtime
Miso TTS is an 8 billion, highly emotive text-to-speech model
Open-source framework for intelligent speech interaction
Multimodal AI Story Teller, built with Stable Diffusion, GPT, etc.