The behavior guidance framework for customer-facing LLM agents
TTS model capable of streaming conversational audio in realtime
Generate audiobooks from e-books
On-device TTS model by Neuphonic
1 min voice data can also be used to train a good TTS model
Instant voice cloning by MIT and MyShell. Audio foundation model
Towards Human-Sounding Speech
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Framework for building realtime multimodal voice AI agents apps
A simple, high-quality voice conversion tool focused on ease of use
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Interface for OuteTTS models
Voice Recognition to Text Tool
A modular voice assistant application for experimenting
Fast multimodal LLM for real-time voice interaction and AI apps
The official Python library for the Fish Audio API
Toolkit for conversational AI
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
A simple native web interface that uses ChatTTS to synthesize text
Generate audiobooks from e-books, voice cloning & 1107+ languages
LLM-based Reinforcement Learning audio edit model
AI-powered tool for generating, optimizing, and translating subtitles
Qwen3-ASR is an open-source series of ASR models
Offline inference engine for art, real-time voice conversations
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD