Minimalistic audiobook player
Comprehensive Gradio WebUI for audio processing
Fast and accurate automatic speech recognition (ASR) for edge devices
Speech to Text to Speech, sends text as OSC messages
GLM-4-Voice | End-to-End Chinese-English Conversational Model
A sound cloning tool with a web interface, using your voice
The open-source voice synthesis studio powered by Qwen3-TTS
Documents and exposes generated source for menu-bar workflows
Instant voice cloning by MIT and MyShell. Audio foundation model
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A simple, high-quality voice conversion tool focused on ease of use
Industrial-level controllable zero-shot text-to-speech system
High-Quality Voice Cloning TTS for 600+ Languages
Telegram Desktop messaging app
In-App assistant SDK to build a multimodal conversational UX websites
1 min voice data can also be used to train a good TTS model
Conversational voice AI agents
On-device wake word detection powered by deep learning
Framework for building real-time voice and multimodal AI agents
AI-polished text appears at your cursor in any app
Official PyTorch Implementation
Realtime AI Voice Agents with SoTA Multimodal AI models on Arduino ESP
Open source AI VTuber platform with voice chat and Live2D avatars
Qwen3-TTS is an open-source series of TTS models
NeuTTS model built from small LLM backbones