CLIP, Predict the most relevant text snippet given an image
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
An experimental version of DeepSeek model
PyTorch code and models for the DINOv2 self-supervised learning
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Unified Multimodal Understanding and Generation Models
FAIR Sequence Modeling Toolkit 2
GPT4V-level open-source multi-modal model based on Llama3-8B
A series of math-specific large language models of our Qwen2 series
Kimi K2 is the large language model series developed by Moonshot AI
RGBD video generation model conditioned on camera input
Generate Any 3D Scene in Seconds
Project Lyra: Open Generative 3D World Models
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
Qwen3-ASR is an open-source series of ASR models
Foundation model for image generation
Block Diffusion for Ultra-Fast Speculative Decoding
Multimodal embedding and reranking models built on Qwen3-VL
Collection of Gemma 3 variants that are trained for performance
Implementation of "MobileCLIP" CVPR 2024
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
Tool for exploring and debugging transformer model behaviors
Qwen2.5-VL is the multimodal large language model series