Speech-to-text, text-to-speech, and speaker recognition
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Long-form streaming TTS system for multi-speaker dialogue generation
Interface for OuteTTS models
Automatic Speech Recognition with Word-level Timestamps
Open-source multi-speaker long-form text-to-speech model
super expressive prompting model based on ltx2.3
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Generate production-ready Lottie animations with Claude Code or Codex
Let the Xiaoai speaker "hear your voice"
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Collaborative document editing using Markdown
A generative speech model for daily dialogue
Sempare Template (scripting) Engine for Delphi
A Web UI for easy subtitle using whisper model
reveal.js on steroids. Get beautiful reveal.js presentations
1B text generation model based on the HRM architecture
Recognition and resolution of numbers, units, date/time, etc.
Buzz transcribes and translates audio offline
MOSS‑TTS Family open‑source speech and sound generation model
An Open Source implementation of Notebook LM with more flexibility
Play ChatGPT and other LLM with Xiaomi AI Speaker
Document (PDF, Word, PPTX ...) extraction and parse API
Oobabooga - The definitive Web UI for local AI, with powerful features
Official PyTorch Implementation