Industrial-level controllable zero-shot text-to-speech system
super expressive prompting model based on ltx2.3
MOSS‑TTS Family open‑source speech and sound generation model
Open-source framework for intelligent speech interaction
Controllable & emotion-expressive zero-shot TTS
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Miso TTS is an 8 billion, highly emotive text-to-speech model
Extension index for stable-diffusion-webui
Long-form streaming TTS system for multi-speaker dialogue generation
A collection of high-quality models for the MuJoCo physics engine
Learning to Act by Watching Unlabeled Online Videos
Dia-1.6B generates lifelike English dialogue and vocal expressions