Making large AI models cheaper, faster and more accessible
Gracefully face hCaptcha challenge with multimodal llms
RF-DETR is a real-time object detection and segmentation
Director, Screenwriter, Producer, and Video Generator All-in-One
The open-source tool for building high-quality datasets
PDF to Markdown with vision models
Reference PyTorch implementation and models for DINOv3
High-performance Inference and Deployment Toolkit for LLMs and VLMs
Open source driver assistance system
An Open Real-time Video-Language Interaction System
Multilingual Document Layout Parsing in a Single Vision-Language Model
A Pragmatic VLA Foundation Model
"Big Model" trains a visual multimodal VLM with 26M parameters
Hub of ready-to-use datasets for ML models
Large-language-model & vision-language-model based on Linear Attention
This repository contains the official implementation of FastVLM
CogView4, CogView3-Plus and CogView3(ECCV 2024)
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
Codex plugin that turns attached object images into code-only
Training data (data labeling, annotation, workflow) for all data types
A framework to enable multimodal models to operate a computer
The largest collection of PyTorch image encoders / backbones
We write your reusable computer vision tools
Free Motion Capture for Everyone
PyTorch code and models for the DINOv2 self-supervised learning