Multi-modal large language model designed for audio understanding
Buzz transcribes and translates audio offline
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Interface for OuteTTS models
VoiceStudio is the open-source, fully-local ElevenLabs alternative
Open-source multi-speaker long-form text-to-speech model
Automatic Speech Recognition with Word-level Timestamps
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A Web UI for easy subtitle using whisper model
Official PyTorch Implementation
Music Assistant is a free, opensource Media library manager
MOSS‑TTS Family open‑source speech and sound generation model
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Googles NotebookLM but local
An Open Source implementation of Notebook LM with more flexibility
Instant voice cloning by MIT and MyShell. Audio foundation model
Translate the video from one language to another and embed dubbing
High-Quality Voice Cloning TTS for 600+ Languages
A Python library for audio
Award-Winning Open Source Video Editing Software
The most powerful and modular diffusion model GUI, api and backend
Speech recognition module for Python
Download videos from websites like YouTube and many others
Edit videos with Claude Code
GenAI Processors is a lightweight Python library