Qwen-Image is a powerful image generation foundation model
Robust Speech Recognition Across Languages, Dialects
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Tiny vision language model
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
scikit-learn compatible tabular foundation model
Model export recipes, Python primitives, and Swift runtime utilities
General-purpose image editing model that delivers high-fidelity
Open-Source Financial Large Language Models
PyTorch code and models for the DINOv2 self-supervised learning
One-click local MCP server installation in desktop apps
A Customizable Image-to-Video Model based on HunyuanVideo
A 0.1B Omni model trained from scratch
26m function call model that runs on incredibly small devices
Video Object and Interaction Deletion
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Phi-3.5 for Mac: Locally-run Vision and Language Models
Qwen2.5-VL is the multimodal large language model series
GLM-4 series: Open Multilingual Multimodal Chat LMs
State-of-the-art (SoTA) text-to-video pre-trained model
HY-Motion model for 3D character animation generation
1B text generation model based on the HRM architecture
Repo for SeedVR2 & SeedVR
Repo of Qwen2-Audio chat & pretrained large audio language model
Tongyi Deep Research, the Leading Open-source Deep Research Agent