A nearly-live implementation of OpenAI's Whisper
Streaming Real-time Audio-Driven Avatar Generation
An Open Source implementation of Notebook LM with more flexibility
SOTA Open Source TTS
Ableton Live Model Context Protocol Integration
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
High-Quality Voice Cloning TTS for 600+ Languages
Converts text to speech in realtime
NetEase cloud music command line version
An open-source, ultra-low-latency remote desktop for Linux hosts
Offline Text To Speech synthesis for python
Generate blog articles from video or audio
AsrTools: Smart Voice-to-Text Tool
Robust Speech Recognition via Large-Scale Weak Supervision
Music Theory for Humans
A youtube-dl fork with additional features and fixes
Use Microsoft Edge's online text-to-speech service from Python
Data Infrastructure providing an approach to multimodal AI workloads
PersonaPlex code
Multimodal-Driven Architecture for Customized Video Generation
Edit videos with Claude Code
Controllable & emotion-expressive zero-shot TTS
Label Studio is a multi-type data labeling and annotation tool
An open-source music player with simple UI
A Web UI for easy subtitle using whisper model