Speech-to-text, text-to-speech, and speaker recognition
Play ChatGPT and other LLM with Xiaomi AI Speaker
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Long-form streaming TTS system for multi-speaker dialogue generation
Interface for OuteTTS models
Automatic Speech Recognition with Word-level Timestamps
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Open-source multi-speaker long-form text-to-speech model
The HTML Presentation Framework
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A Web UI for easy subtitle using whisper model
A LaTeX class for producing presentations and slides
super expressive prompting model based on ltx2.3
OpenVoiceOS Core, the FOSS Artificial Intelligence platform
An Open Source implementation of Notebook LM with more flexibility
Official PyTorch Implementation
Ultraminimalist macOS recording + transcription
The ioquake3 community effort to continue supporting/developing id's
Self-hosted AI audio transcription
The official KotlinConf application
Control SONOS speakers from your terminal
A private, local meeting notes assistant
macOS System-wide audio equalizer & volume mixer
MOSS‑TTS Family open‑source speech and sound generation model
Translate the video from one language to another and embed dubbing