Multi-modal large language model designed for audio understanding
Speech-to-text, text-to-speech, and speaker recognition
Audio Share can share Windows/Linux computer's audio to Android phone
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Interface for OuteTTS models
Open-source multi-speaker long-form text-to-speech model
Automatic Speech Recognition with Word-level Timestamps
Clone a voice in 5 seconds to generate arbitrary speech in real-time
super expressive prompting model based on ltx2.3
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
A Web UI for easy subtitle using whisper model
Self-hosted AI audio transcription
A private, local meeting notes assistant
Official PyTorch Implementation
An Open Source implementation of Notebook LM with more flexibility
The ioquake3 community effort to continue supporting/developing id's
Music Assistant is a free, opensource Media library manager
MOSS‑TTS Family open‑source speech and sound generation model
Open source software for live streaming and recording
Control SONOS speakers from your terminal
SpatGRIS4
The HTML Presentation Framework
Instantly generate AI-powered subtitles on your device
Googles NotebookLM but local
Audio Plugin for Audio to MIDI transcription using deep learning