Controllable and fast Text-to-Speech for over 7000 languages
Real-time voice interactive digital human
A lightweight text-to-speech model with zero-shot voice cloning
StreamSpeech is a seamless model for offline speech recognition
Toolkit for conversational AI
Synchronized Translation for Videos
Clone a voice in 5 seconds to generate arbitrary speech in real-time