Qwen-Image is a powerful image generation foundation model
Foundation model for image generation
Qwen's most powerful open-source image generation model
General-purpose image editing model that delivers high-fidelity
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
CLIP, Predict the most relevant text snippet given an image
Official inference repo for FLUX.1 models
Official inference repo for FLUX.2 models
A Powerful Native Multimodal Model for Image Generation
Open image model at the forefront of design
Wan2.1: Open and Advanced Large-Scale Video Generative Model
MiniMax H3 is a general-purpose, omni-modal generative system
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A Unified Framework for Text-to-3D and Image-to-3D Generation
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Open-source image generative foundation model
Collection of Gemma 3 variants that are trained for performance
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Text and image to video generation: CogVideoX and CogVideo
Multimodal-Driven Architecture for Customized Video Generation
Generating Immersive, Explorable, and Interactive 3D Worlds
Contexts Optical Compression
Capable of understanding text, audio, vision, video
AI PPT Track Terminator, the strongest PPT Skill ever
CogView4, CogView3-Plus and CogView3(ECCV 2024)