ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Buzz transcribes and translates audio offline
Play ChatGPT and other LLM with Xiaomi AI Speaker
VoiceStudio is the open-source, fully-local ElevenLabs alternative
Long-form streaming TTS system for multi-speaker dialogue generation
Automatic Speech Recognition with Word-level Timestamps
Interface for OuteTTS models
OpenVoiceOS Core, the FOSS Artificial Intelligence platform
Open-source multi-speaker long-form text-to-speech model
Official PyTorch Implementation
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A Web UI for easy subtitle using whisper model
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
An Open Source implementation of Notebook LM with more flexibility
A PyTorch-based Speech Toolkit
MOSS‑TTS Family open‑source speech and sound generation model
A generative speech model for daily dialogue
Translate the video from one language to another and embed dubbing
Googles NotebookLM but local
Music Assistant is a free, opensource Media library manager
Multi-modal large language model designed for audio understanding
High-Quality Voice Cloning TTS for 600+ Languages
Instant voice cloning by MIT and MyShell. Audio foundation model
Robust Speech Recognition Across Languages, Dialects
Spark-TTS Inference Code