Multi-modal large language model designed for audio understanding
Speech-to-text, text-to-speech, and speaker recognition
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Interface for OuteTTS models
Open-source multi-speaker long-form text-to-speech model
Automatic Speech Recognition with Word-level Timestamps
super expressive prompting model based on ltx2.3
Clone a voice in 5 seconds to generate arbitrary speech in real-time
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
A Web UI for easy subtitle using whisper model
Self-hosted AI audio transcription
A private, local meeting notes assistant
Official PyTorch Implementation
An Open Source implementation of Notebook LM with more flexibility
MOSS‑TTS Family open‑source speech and sound generation model
Googles NotebookLM but local
Instantly generate AI-powered subtitles on your device
Audio Plugin for Audio to MIDI transcription using deep learning
High-Quality Voice Cloning TTS for 600+ Languages
Instant voice cloning by MIT and MyShell. Audio foundation model
Web presentation editor replicating many PowerPoint features online
The open-source voice synthesis studio powered by Qwen3-TTS
The most powerful and modular diffusion model GUI, api and backend
A Python library for audio
Speech recognition module for Python