A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Official MiniMax Model Context Protocol (MCP) server
Comprehensive Gradio WebUI for audio processing
Real-time voice interactive digital human
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Miso TTS is an 8 billion, highly emotive text-to-speech model
SOTA Open Source TTS
AI-Powered Wiki Generator for GitHub/Gitlab/Bitbucket Repositories
Long-form streaming TTS system for multi-speaker dialogue generation
An Open Source text-to-speech system built by inverting Whisper
Towards Human-Sounding Speech
Foundational model for human-like, expressive TTS
MARS5 speech model (TTS) from CAMB.AI
A deep learning toolkit for Text-to-Speech, battle-tested in research
Multi-Voice and Prompt-Controlled TTS Engine
Open source implementation of Microsoft's VALL-E X zero-shot TTS model
Simple and powerful voice changer for Linux, written with Python & GTK
A webui for different audio related Neural Networks
Singing voice change based on whisper, lora for singing voice clone
A collection of high-quality models for the MuJoCo physics engine
Learning to Act by Watching Unlabeled Online Videos
[WIP] VoiceSmith makes training text to speech models easy
A Python/Pytorch app for easily synthesising human voices
A library of additional estimators and SageMaker tools based on scikit