Diffusion Transformer with Fine-Grained Chinese Understanding
Fast stable diffusion on CPU and AI PC
Qwen3-omni is a natively end-to-end, omni-modal LLM
Code for running inference and finetuning with SAM 3 model
A 0.1B Omni model trained from scratch
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
A Multi-Modal World Model for Reconstructing, Generating, Simulation
High-Resolution Image Synthesis with Latent Diffusion Models
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
Accurate × Fast × Comprehensive
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Unified Multimodal Understanding and Generation Models
Fast-stable-diffusion + DreamBooth
Implementation of "MobileCLIP" CVPR 2024
Multimodal embedding and reranking models built on Qwen3-VL
Generate Any 3D Scene in Seconds
Official Python inference and LoRA trainer package
Phi-3.5 for Mac: Locally-run Vision and Language Models
Large-language-model & vision-language-model based on Linear Attention
Official implementation of DreamCraft3D
A Systematic Framework for Interactive World Modeling
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Infinite Worlds with Versatile Interactions
Chinese and English multimodal conversational language model