GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
GLM-4 series: Open Multilingual Multimodal Chat LMs
Diversity-driven optimization and large-model reasoning ability
This repository contains the official implementation of FastVLM
PyTorch code and models for the DINOv2 self-supervised learning
CogView4, CogView3-Plus and CogView3(ECCV 2024)
A Multi-Modal World Model for Reconstructing, Generating, Simulation
code for Mesh R-CNN, ICCV 2019
DeepSeek Coder: Let the Code Write Itself
Designed for text embedding and ranking tasks
Codex plugin that turns attached object images into code-only
Audio Language Models are Few-Shot Learners
Open-source industrial-grade ASR models
Foundation model for image generation
Block Diffusion for Ultra-Fast Speculative Decoding
Multimodal embedding and reranking models built on Qwen3-VL
Implementation of "MobileCLIP" CVPR 2024
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
Video understanding codebase from FAIR for reproducing video models
Tool for exploring and debugging transformer model behaviors
CLIP, Predict the most relevant text snippet given an image
Ling is a MoE LLM provided and open-sourced by InclusionAI
Tongyi Deep Research, the Leading Open-source Deep Research Agent
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning