Industrial-level controllable zero-shot text-to-speech system
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Towards Human-Sounding Speech
StreamSpeech is a seamless model for offline speech recognition
A TTS model capable of generating ultra-realistic dialogue
A lightweight text-to-speech model with zero-shot voice cloning
A high-quality rapid TTS voice cloning model
Miso TTS is an 8 billion, highly emotive text-to-speech model
Scalable generative AI framework built for researchers and developers
Interface for OuteTTS models
Spark-TTS Inference Code
A text-to-speech, speech-to-text and speech-to-speech library
Automatically translates the text of a video based on a subtitle file
Foundational model for human-like, expressive TTS
Framework for building neural networks
Controllable & emotion-expressive zero-shot TTS
Management of Yandex Station and other smart home devices
Controllable and fast Text-to-Speech for over 7000 languages
Build Vision Agents quickly with any model or video provider
Reading book source
Python library and CLI tool to interface with Google Translate
Toolkit for conversational AI
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
One-click deployment (including offline integration package)