An Open Real-time Video-Language Interaction System
A Pragmatic VLA Foundation Model
ICLR2024 Spotlight: curation/training code, metadata, distribution
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
Reference PyTorch implementation and models for DINOv3
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
PyTorch code and models for the DINOv2 self-supervised learning
Qwen2.5-VL is the multimodal large language model series
OCR expert VLM powered by Hunyuan's native multimodal architecture
Contexts Optical Compression
Tooling for the Common Objects In 3D dataset
Official DeiT repository
PyTorch implementation of MAE