Code for running inference and finetuning with SAM 3 model
Phi-3.5 for Mac: Locally-run Vision and Language Models
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
A 0.1B Omni model trained from scratch
Implementation of "MobileCLIP" CVPR 2024
Accurate × Fast × Comprehensive
Official implementation of DreamCraft3D
RGBD video generation model conditioned on camera input
Infinite Worlds with Versatile Interactions
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Qwen3-omni is a natively end-to-end, omni-modal LLM
Advancing Open-source World Models
Generate Any 3D Scene in Seconds
PyTorch code and models for the DINOv2 self-supervised learning
A Systematic Framework for Interactive World Modeling
Project Lyra: Open Generative 3D World Models
Multimodal embedding and reranking models built on Qwen3-VL
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
code for Mesh R-CNN, ICCV 2019
Language modeling in a sentence representation space
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Large-language-model & vision-language-model based on Linear Attention
Chinese and English multimodal conversational language model
High-Resolution Image Synthesis with Latent Diffusion Models