Repo of Qwen2-Audio chat & pretrained large audio language model
Open-source large language model family from Tencent Hunyuan
Tool for exploring and debugging transformer model behaviors
Implementation of "MobileCLIP" CVPR 2024
PyTorch code and models for the DINOv2 self-supervised learning
Qwen2.5-VL is the multimodal large language model series
An Open Real-time Video-Language Interaction System
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Generating Immersive, Explorable, and Interactive 3D Worlds
A Unified Framework for Text-to-3D and Image-to-3D Generation
Tongyi Deep Research, the Leading Open-source Deep Research Agent
AI cognitive-enhancement Skills based on Anthropic's J-space
Qwen3-ASR is an open-source series of ASR models
New family of code large language models (LLMs)
Visual Causal Flow
Inference code for scalable emulation of protein equilibrium ensembles
Video Object and Interaction Deletion
VMZ: Model Zoo for Video Modeling
Python SDK for Claude Agent
CLIP, Predict the most relevant text snippet given an image
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
4M: Massively Multimodal Masked Modeling
Hackable and optimized Transformers building blocks
Official implementation of DreamCraft3D