Play ChatGPT and other LLM with Xiaomi AI Speaker
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Long-form streaming TTS system for multi-speaker dialogue generation
Interface for OuteTTS models
Automatic Speech Recognition with Word-level Timestamps
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Open-source multi-speaker long-form text-to-speech model
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A Web UI for easy subtitle using whisper model
An Open Source implementation of Notebook LM with more flexibility
OpenVoiceOS Core, the FOSS Artificial Intelligence platform
Official PyTorch Implementation
MOSS‑TTS Family open‑source speech and sound generation model
Translate the video from one language to another and embed dubbing
A PyTorch-based Speech Toolkit
High-Quality Voice Cloning TTS for 600+ Languages
A generative speech model for daily dialogue
Multi-modal large language model designed for audio understanding
Music Assistant is a free, opensource Media library manager
Instant voice cloning by MIT and MyShell. Audio foundation model
Googles NotebookLM but local
Spark-TTS Inference Code
Robust Speech Recognition Across Languages, Dialects
End-to-end speech processing toolkit
Foundational model for human-like, expressive TTS