Multi-modal large language model designed for audio understanding
Speech-to-text, text-to-speech, and speaker recognition
Audio Share can share Windows/Linux computer's audio to Android phone
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Interface for OuteTTS models
Open-source multi-speaker long-form text-to-speech model
Ultraminimalist macOS recording + transcription
Automatic Speech Recognition with Word-level Timestamps
macOS System-wide audio equalizer & volume mixer
super expressive prompting model based on ltx2.3
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A native macOS menu bar app for managing audio device priorities
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
A Web UI for easy subtitle using whisper model
Self-hosted AI audio transcription
A private, local meeting notes assistant
Official PyTorch Implementation
An Open Source implementation of Notebook LM with more flexibility
High-performance audio engine for react-native
Music Assistant is a free, opensource Media library manager
Open source software for live streaming and recording
MOSS‑TTS Family open‑source speech and sound generation model
The ioquake3 community effort to continue supporting/developing id's
Control SONOS speakers from your terminal
The HTML Presentation Framework