Advanced techniques for RAG systems
Omnilingual ASR Open-Source Multilingual SpeechRecognition
Fast and Universal 3D reconstruction model for versatile tasks
Implementation of Vision Transformer, a simple way to achieve SOTA
4M: Massively Multimodal Masked Modeling
Refer and Ground Anything Anywhere at Any Granularity
A Model Context Protocol (MCP) Gateway & Registry
Set of tools to assess and improve LLM security
Agent toolkit providing semantic retrieval and editing capabilities
Open-source platform for building enterprise-grade agents
ICLR2024 Spotlight: curation/training code, metadata, distribution
PyTorch code and models for V-JEPA self-supervised learning from video
A PyTorch library for implementing flow matching algorithms
An implementation of a deep learning recommendation model (DLRM)
Self-supervised visual learning using momentum contrast in PyTorch
PyTorch3D is FAIR's library of reusable components for deep learning
[CVPR 2025 Best Paper Award] VGGT
Anthropic's educational courses
Official implementation of DreamCraft3D
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
Research code artifacts for Code World Model (CWM)
Diffusion Transformer with Fine-Grained Chinese Understanding
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
The Memory layer for AI Agents