Capable of understanding text, audio, vision, video
Multimodal-Driven Architecture for Customized Video Generation
TTS model capable of streaming conversational audio in realtime
Controllable & emotion-expressive zero-shot TTS
Generate audiobooks from e-books, voice cloning & 1107+ languages
Open-source multi-speaker long-form text-to-speech model
Speech recognition module for Python
MOSS-TTS-Nano is an open-source multilingual tiny speech generation
Official Python inference and LoRA trainer package
Industrial-level controllable zero-shot text-to-speech system
Instant voice cloning by MIT and MyShell. Audio foundation model
Automatic subtitle synchronization tool
A nearly-live implementation of OpenAI's Whisper
Speakr is a personal, self-hosted web application
Transforming Multimodal Content into Captivating Multilingual Audio
A speech-text foundation model for real time dialogue
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Free, high-quality text-to-speech API endpoint to replace OpenAI
Generate audiobooks from e-books
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Offline Text To Speech synthesis for python
A lightweight text-to-speech model with zero-shot voice cloning
Fast multimodal LLM for real-time voice interaction and AI apps
Interface for OuteTTS models
PersonaPlex code