code for Mesh R-CNN, ICCV 2019
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
High-Fidelity and Controllable Generation of Textured 3D Assets
Bidirectional token-classification model for identifiable info
Genome modeling and design across all domains of life
Project Lyra: Open Generative 3D World Models
Diffusion Transformer with Fine-Grained Chinese Understanding
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Tiny vision language model
Repo for SeedVR2 & SeedVR
A Multi-Modal World Model for Reconstructing, Generating, Simulation
DeepMind model for tracking arbitrary points across videos & robotics
Tooling for the Common Objects In 3D dataset
Uncommon Objects in 3D dataset
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
GPT4V-level open-source multi-modal model based on Llama3-8B
Designed for text embedding and ranking tasks
Generating Immersive, Explorable, and Interactive 3D Worlds
State-of-the-art (SoTA) text-to-video pre-trained model
LLM-based Reinforcement Learning audio edit model
Audio Language Models are Few-Shot Learners