State-of-the-art TTS model under 25MB
MOSS‑TTS Family open‑source speech and sound generation model
Framework for building real-time voice and multimodal AI agents
ImageBind One Embedding Space to Bind Them All
A modular voice assistant application for experimenting
A 0.1B Omni model trained from scratch
Unified web UI for training and running open models locally
Collect your thoughts and notes without leaving the command line
An Open Source text-to-speech system built by inverting Whisper
Translate the video from one language to another and embed dubbing
Strip multi-vendor AI provenance marks
StreamSpeech is a seamless model for offline speech recognition
AI-powered tool for generating, optimizing, and translating subtitles
Pushing the Frontier of Long Audio-Visual Generation
Multimodal Diffusion with Representation Alignment
Implementation of AudioLM audio generation model in Pytorch
Foundational video generation model with 13.6B parameters
GLM-4-Voice | End-to-End Chinese-English Conversational Model
GenAI Processors is a lightweight Python library
Multi-lingual large voice generation model, providing inference
Framework for building realtime multimodal voice AI agents apps
Qwen3-ASR is an open-source series of ASR models
Open source AI model for generating full songs from lyrics prompts
Official MiniMax Model Context Protocol (MCP) server
A sound cloning tool with a web interface, using your voice