RGBD video generation model conditioned on camera input
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Official Python inference and LoRA trainer package
Sharp Monocular Metric Depth in Less Than a Second
Generate Any 3D Scene in Seconds
Metric monocular depth estimation (vision model)
Open VLA model for autonomous driving reasoning and planning
Vision-language-action model for robot control via images and text