Speech-to-text, text-to-speech, and speaker recognition
Let the Xiaoai speaker "hear your voice"
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Buzz transcribes and translates audio offline
Play ChatGPT and other LLM with Xiaomi AI Speaker
VoiceStudio is the open-source, fully-local ElevenLabs alternative
Long-form streaming TTS system for multi-speaker dialogue generation
Automatic Speech Recognition with Word-level Timestamps
Interface for OuteTTS models
OpenVoiceOS Core, the FOSS Artificial Intelligence platform
The HTML Presentation Framework
reveal.js on steroids. Get beautiful reveal.js presentations
Open-source multi-speaker long-form text-to-speech model
Let the Xiaoai speaker "listen to you"
A LaTeX class for producing presentations and slides
Official PyTorch Implementation
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A Web UI for easy subtitle using whisper model
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
An Open Source implementation of Notebook LM with more flexibility
The ioquake3 community effort to continue supporting/developing id's
super expressive prompting model based on ltx2.3
A PyTorch-based Speech Toolkit
Control SONOS speakers from your terminal
Self-hosted AI audio transcription