State-of-the-art (SoTA) text-to-video pre-trained model
Lets make video diffusion practical
Video understanding codebase from FAIR for reproducing video models
MiniMax H3 is a general-purpose, omni-modal generative system
Official Python inference and LoRA trainer package
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
MiniMax H3 inference engine for Mac computers
Repo for SeedVR2 & SeedVR
Multimodal-Driven Architecture for Customized Video Generation
Video Object and Interaction Deletion
RGBD video generation model conditioned on camera input
Inference script for Oasis 500M
Uncommon Objects in 3D dataset
Project Lyra: Open Generative 3D World Models
OCR expert VLM powered by Hunyuan's native multimodal architecture
The official pytorch implementation of our paper
Google’s flagship dense multimodal model for coding and reasoning