Agentic, Reasoning, and Coding (ARC) foundation models
Code for running inference and finetuning with SAM 3 model
Advanced language and coding AI model
Implementation of "MobileCLIP" CVPR 2024
PyTorch code and models for the DINOv2 self-supervised learning
CLIP, Predict the most relevant text snippet given an image
Large Multimodal Models for Video Understanding and Editing
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Open Multilingual Multimodal Chat LMs
The ChatGPT Retrieval Plugin lets you easily find personal documents
Reference implementation of the Transformer architecture optimized