Instant voice cloning by MIT and MyShell. Audio foundation model
A Family of Open Sourced Music Foundation Models
Interface for OuteTTS models
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
An open-source toolkit for BigMac-style pipeline-parallel training
A lightweight text-to-speech model with zero-shot voice cloning
SOTA Open Source TTS
Taming Stable Diffusion for Lip Sync
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
Multi-lingual large voice generation model, providing inference
Run PyTorch LLMs locally on servers, desktop and mobile
Official code for Style Aligned Image Generation via Shared Attention
Open source implementation of Microsoft's VALL-E X zero-shot TTS model
Clone a voice in 5 seconds to generate arbitrary speech in real-time
The basic distribution probability Tutorial for Deep Learning Research