Accurate × Fast × Comprehensive
Miso TTS is an 8 billion, highly emotive text-to-speech model
Qwen3-omni is a natively end-to-end, omni-modal LLM
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Capable of understanding text, audio, vision, video
Open-source image generative foundation model
Multimodal-Driven Architecture for Customized Video Generation
Designed for text embedding and ranking tasks
Generating Immersive, Explorable, and Interactive 3D Worlds
Collection of Gemma 3 variants that are trained for performance
Long-form streaming TTS system for multi-speaker dialogue generation
tiktoken is a fast BPE tokeniser for use with OpenAI's models
State-of-the-art (SoTA) text-to-video pre-trained model
Foundation model for image generation
The most powerful local music generation model
General-purpose image editing model that delivers high-fidelity
A 0.1B Omni model trained from scratch
Open Source Speech Language Model
Multimodal embedding and reranking models built on Qwen3-VL
Official Python inference and LoRA trainer package
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Qwen2.5-VL is the multimodal large language model series
Fast stable diffusion on CPU and AI PC
HY-Motion model for 3D character animation generation
Large-language-model & vision-language-model based on Linear Attention