TTS with kokoro and onnx runtime
Industrial-level controllable zero-shot text-to-speech system
Framework for building neural networks
Automatically translates the text of a video based on a subtitle file
Generate audiobooks from e-books
End-to-end speech processing toolkit
A nearly-live implementation of OpenAI's Whisper
A TTS model capable of generating ultra-realistic dialogue
Build Vision Agents quickly with any model or video provider
A simple native web interface that uses ChatTTS to synthesize text
Controllable & emotion-expressive zero-shot TTS
SOTA discrete acoustic codec models with 40/75 tokens per second
Official MiniMax Model Context Protocol (MCP) server
A fast TTS architecture with conditional flow matching
A sound cloning tool with a web interface, using your voice
Offline desktop app to convert EPUB to MP3 using Kokoro-82M neural TTS
VITS2 backbone with multilingual-bert
Unofficial Parallel WaveGAN
A webui for different audio related Neural Networks
WaveRNN Vocoder + TTS
General Speech Restoration
Real-Time State-of-the-art Speech Synthesis for Tensorflow 2
Implementation of a Transformer based neural network
Conditional Variational Autoencoder with Adversarial Learning
Generative Adversarial Networks for Efficient and High Fidelity Speech