Multi-modal large language model designed for audio understanding
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Interface for OuteTTS models
Open-source multi-speaker long-form text-to-speech model
Automatic Speech Recognition with Word-level Timestamps
Clone a voice in 5 seconds to generate arbitrary speech in real-time
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
A Web UI for easy subtitle using whisper model
Official PyTorch Implementation
An Open Source implementation of Notebook LM with more flexibility
Music Assistant is a free, opensource Media library manager
MOSS‑TTS Family open‑source speech and sound generation model
Googles NotebookLM but local
High-Quality Voice Cloning TTS for 600+ Languages
Instant voice cloning by MIT and MyShell. Audio foundation model
Translate the video from one language to another and embed dubbing
The most powerful and modular diffusion model GUI, api and backend
A Python library for audio
Award-Winning Open Source Video Editing Software
Speech recognition module for Python
Mopidy is an extensible music server written in Python
Download videos from websites like YouTube and many others
GenAI Processors is a lightweight Python library
Video editing with Python
pyglet is a cross-platform windowing and multimedia library for Python