Foundational video generation model with 13.6B parameters
Give Claude the ability to watch any video
State-of-the-art (SoTA) text-to-video pre-trained model
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
Let Claude (or any LLM) actually watch a video
Lets make video diffusion practical
AI tool that removes hardcoded subtitles and text from videos locally
Taming Stable Diffusion for Lip Sync
Video understanding codebase from FAIR for reproducing video models
Real time face swap and one-click video deepfake
Video-based AI memory library. Store millions of text chunks in MP4
MiniMax H3 is a general-purpose, omni-modal generative system
Official Python inference and LoRA trainer package
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
Official Repo For "Sa2VA: Marrying SAM2 with LLaVA
Repo for SeedVR2 & SeedVR
Multimodal-Driven Architecture for Customized Video Generation
Build Vision Agents quickly with any model or video provider
Video Object and Interaction Deletion
RGBD video generation model conditioned on camera input
Inference script for Oasis 500M
Uncommon Objects in 3D dataset
Pluggable SOTA multi-object tracking modules for segmentation
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
Project Lyra: Open Generative 3D World Models