A generative speech model for daily dialogue
Open-source framework for intelligent speech interaction
Controllable & emotion-expressive zero-shot TTS
Towards Human-Sounding Speech
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
A fast TTS architecture with conditional flow matching
Multimodal AI Story Teller, built with Stable Diffusion, GPT, etc.
PyTorch implementation of VALL-E (Zero-Shot Text-To-Speech)