ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Open Source Speech Language Model
SOTA discrete acoustic codec models with 40/75 tokens per second
A high-quality rapid TTS voice cloning model
Robust Speech Recognition via Large-Scale Weak Supervision
Qwen3-TTS is an open-source series of TTS models
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
The official Python SDK for the ElevenLabs API
Code and models for ICML 2024 paper, NExT-GPT
Sample code and notebooks for Generative AI on Google Cloud
A Web UI for easy subtitle using whisper model
ImageBind One Embedding Space to Bind Them All
Voice Recognition to Text Tool
Pushing the Frontier of Long Audio-Visual Generation
Python library and CLI tool to interface with Google Translate
A simple native web interface that uses ChatTTS to synthesize text
An Open Source implementation of Notebook LM with more flexibility
Towards Human-Sounding Speech
State-of-the-art TTS model under 25MB
MOSS‑TTS Family open‑source speech and sound generation model
Framework for building real-time voice and multimodal AI agents
Unified web UI for training and running open models locally
Build Vision Agents quickly with any model or video provider
A modular voice assistant application for experimenting