This repository contains the official implementation of FastVLM
PyTorch code and models for the DINOv2 self-supervised learning
CogView4, CogView3-Plus and CogView3(ECCV 2024)
code for Mesh R-CNN, ICCV 2019
DeepSeek Coder: Let the Code Write Itself
Generating Immersive, Explorable, and Interactive 3D Worlds
Diversity-driven optimization and large-model reasoning ability
Repo of Qwen2-Audio chat & pretrained large audio language model
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Audio Language Models are Few-Shot Learners
Open-source industrial-grade ASR models
Foundation model for image generation
Block Diffusion for Ultra-Fast Speculative Decoding
Multimodal embedding and reranking models built on Qwen3-VL
Implementation of "MobileCLIP" CVPR 2024
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
Tool for exploring and debugging transformer model behaviors
CLIP, Predict the most relevant text snippet given an image
Ling is a MoE LLM provided and open-sourced by InclusionAI
Qwen2.5-VL is the multimodal large language model series
A Multi-Modal World Model for Reconstructing, Generating, Simulation
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A series of math-specific large language models of our Qwen2 series