A 0.1B Omni model trained from scratch
An Open Source text-to-speech system built by inverting Whisper
StreamSpeech is a seamless model for offline speech recognition
Multimodal Diffusion with Representation Alignment
Foundational video generation model with 13.6B parameters
Implementation of AudioLM audio generation model in Pytorch
AI-powered tool for generating, optimizing, and translating subtitles
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Multi-lingual large voice generation model, providing inference
Open source AI model for generating full songs from lyrics prompts
Qwen3-ASR is an open-source series of ASR models
Framework for building realtime multimodal voice AI agents apps
Official MiniMax Model Context Protocol (MCP) server
Edit videos with Claude Code
A sound cloning tool with a web interface, using your voice
A Systematic Framework for Interactive World Modeling
State-of-the-art diffusion models for image and audio generation
Marrying Grounding DINO with Segment Anything & Stable Diffusion
Data Infrastructure providing an approach to multimodal AI workloads
Build multimodal language agents for fast prototype and production
Controllable and fast Text-to-Speech for over 7000 languages
A fast TTS architecture with conditional flow matching
Make any agent harness multimodal-native
Easy-to-use Speech Toolkit including Self-Supervised Learning model
A python tool that uses GPT-4, FFmpeg, and OpenCV