Speech-to-text, text-to-speech, and speaker recognition
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Long-form streaming TTS system for multi-speaker dialogue generation
Interface for OuteTTS models
Automatic Speech Recognition with Word-level Timestamps
Open-source multi-speaker long-form text-to-speech model
super expressive prompting model based on ltx2.3
Clone a voice in 5 seconds to generate arbitrary speech in real-time
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Let the Xiaoai speaker "hear your voice"
Generate production-ready Lottie animations with Claude Code or Codex
A generative speech model for daily dialogue
Collaborative document editing using Markdown
A Web UI for easy subtitle using whisper model
reveal.js on steroids. Get beautiful reveal.js presentations
Sempare Template (scripting) Engine for Delphi
1B text generation model based on the HRM architecture
Recognition and resolution of numbers, units, date/time, etc.
Buzz transcribes and translates audio offline
MOSS‑TTS Family open‑source speech and sound generation model
An Open Source implementation of Notebook LM with more flexibility
Play ChatGPT and other LLM with Xiaomi AI Speaker
Document (PDF, Word, PPTX ...) extraction and parse API
Official PyTorch Implementation
Ultraminimalist macOS recording + transcription