Implementation of Vision Transformer, a simple way to achieve SOTA
4M: Massively Multimodal Masked Modeling
Refer and Ground Anything Anywhere at Any Granularity
Supercharge Your LLM with the Fastest KV Cache Layer
A Model Context Protocol (MCP) Gateway & Registry
Open-source platform for building enterprise-grade agents
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
MobileLLM Optimizing Sub-billion Parameter Language Models
ICLR2024 Spotlight: curation/training code, metadata, distribution
A Production-ready Reinforcement Learning AI Agent Library
PyTorch code and models for V-JEPA self-supervised learning from video
A PyTorch library for implementing flow matching algorithms
An implementation of a deep learning recommendation model (DLRM)
Self-supervised visual learning using momentum contrast in PyTorch
ImageBind One Embedding Space to Bind Them All
[CVPR 2025 Best Paper Award] VGGT
Official implementation of DreamCraft3D
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
Research code artifacts for Code World Model (CWM)
Diffusion Transformer with Fine-Grained Chinese Understanding
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
The Memory layer for AI Agents
A simple screen parsing tool towards pure vision based GUI agent
Low-latency REST API for serving text-embeddings