MiniMax H3 is a general-purpose, omni-modal generative system
AI Fully Automated Short Video Engine
Let Claude (or any LLM) actually watch a video
Repo for SeedVR2 & SeedVR
Inference script for Oasis 500M
AI logo animation skill: turn raster logos into smooth SVG animation
Agent Skill for generating 2D sprite sheets and map, transparent PNG
AI tool that removes hardcoded subtitles and text from videos locally
NVR with realtime local object detection for IP cameras
Official Python inference and LoRA trainer package
Lets make video diffusion practical
Multimodal Diffusion with Representation Alignment
ComfyUI wrapper nodes for WanVideo and related models
A Customizable Image-to-Video Model based on HunyuanVideo
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
Official repository for LTX-Video
Uncommon Objects in 3D dataset
Give Claude the ability to watch any video
Open-source multi-speaker long-form text-to-speech model
Inference code for scalable emulation of protein equilibrium ensembles
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Convert AI papers to GUI
Code for running inference with the SAM 3D Body Model 3DB
Video understanding codebase from FAIR for reproducing video models
Qwen2.5-VL is the multimodal large language model series