Speech-to-text, text-to-speech, and speaker recognition
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Interface for OuteTTS models
Open-source multi-speaker long-form text-to-speech model
Automatic Speech Recognition with Word-level Timestamps
Clone a voice in 5 seconds to generate arbitrary speech in real-time
super expressive prompting model based on ltx2.3
A Web UI for easy subtitle using whisper model
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Self-hosted AI audio transcription
A private, local meeting notes assistant
An Open Source implementation of Notebook LM with more flexibility
MOSS‑TTS Family open‑source speech and sound generation model
Googles NotebookLM but local
Instantly generate AI-powered subtitles on your device
Audio Plugin for Audio to MIDI transcription using deep learning
Instant voice cloning by MIT and MyShell. Audio foundation model
High-Quality Voice Cloning TTS for 600+ Languages
Web presentation editor replicating many PowerPoint features online
The open-source voice synthesis studio powered by Qwen3-TTS
The most powerful and modular diffusion model GUI, api and backend
A Python library for audio
Captcha solver extension for humans
Generate music based on natural language prompts using LLMs