Accurate × Fast × Comprehensive
A Unified Framework for Text-to-3D and Image-to-3D Generation
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
GLM-4-Voice | End-to-End Chinese-English Conversational Model
AI PPT Track Terminator, the strongest PPT Skill ever
State-of-the-art (SoTA) text-to-video pre-trained model
Official Python inference and LoRA trainer package
Controllable & emotion-expressive zero-shot TTS
Long-form streaming TTS system for multi-speaker dialogue generation
The most powerful local music generation model
Miso TTS is an 8 billion, highly emotive text-to-speech model
A 0.1B Omni model trained from scratch
Open Source Speech Language Model
Multimodal embedding and reranking models built on Qwen3-VL
Multimodal-Driven Architecture for Customized Video Generation
Foundation model for image generation
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A Multi-Modal World Model for Reconstructing, Generating, Simulation
FAIR Sequence Modeling Toolkit 2
Capable of understanding text, audio, vision, video
High-Resolution Image Synthesis with Latent Diffusion Models
Qwen3-ASR is an open-source series of ASR models
Qwen2.5-VL is the multimodal large language model series
Fast stable diffusion on CPU and AI PC
Generate Any 3D Scene in Seconds