MiniMax H3 is a general-purpose, omni-modal generative system
AI Fully Automated Short Video Engine
Repo for SeedVR2 & SeedVR
Inference script for Oasis 500M
AI logo animation skill: turn raster logos into smooth SVG animation
Let Claude (or any LLM) actually watch a video
AI tool that removes hardcoded subtitles and text from videos locally
Agent Skill for generating 2D sprite sheets and map, transparent PNG
NVR with realtime local object detection for IP cameras
Official Python inference and LoRA trainer package
ComfyUI wrapper nodes for WanVideo and related models
A Customizable Image-to-Video Model based on HunyuanVideo
Lets make video diffusion practical
Multimodal Diffusion with Representation Alignment
Official repository for LTX-Video
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
Uncommon Objects in 3D dataset
Give Claude the ability to watch any video
Open-source multi-speaker long-form text-to-speech model
Code for running inference with the SAM 3D Body Model 3DB
Inference code for scalable emulation of protein equilibrium ensembles
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Qwen2.5-VL is the multimodal large language model series
Video understanding codebase from FAIR for reproducing video models
A reactive notebook for Python