SoTA open-source TTS
Capable of understanding text, audio, vision, video
Spark-TTS Inference Code
kaldi-asr/kaldi is the official location of the Kaldi project
Self-host the powerful Chatterbox TTS model
VoiceStudio is the open-source, fully-local ElevenLabs alternative
A robust, efficient, low-latency speech-to-text library
Speakr is a personal, self-hosted web application
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Free, high-quality text-to-speech API endpoint to replace OpenAI
NeuTTS model built from small LLM backbones
On-device TTS model by Neuphonic
Official PyTorch Implementation
Instant voice cloning by MIT and MyShell. Audio foundation model
Open-source industrial-grade ASR models
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Framework for building real-time voice and multimodal AI agents
An Open Source text-to-speech system built by inverting Whisper
Generate audiobooks from e-books
Interface for OuteTTS models
A nearly-live implementation of OpenAI's Whisper
TTS model capable of streaming conversational audio in realtime
LLM Large Model of Selling Anchor
SOTA discrete acoustic codec models with 40/75 tokens per second
Multi-lingual large voice generation model, providing inference