MiniMax H3 is a general-purpose, omni-modal generative system
Make videos programmatically with React
Inference script for Oasis 500M
MiniMax H3 inference engine for Mac computers
NVR with realtime local object detection for IP cameras
Official Python inference and LoRA trainer package
Give Claude the ability to watch and understand videos
Multimodal Diffusion with Representation Alignment
A Customizable Image-to-Video Model based on HunyuanVideo
Official repository for LTX-Video
Open-source multi-speaker long-form text-to-speech model
Behavior tree AI for Godot Engine
Code for running inference with the SAM 3D Body Model 3DB
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
AI-assisted storyboard and video generation tool
Video understanding codebase from FAIR for reproducing video models
MCP server enabling AI coding tools to access Figma design data
Large Multimodal Models for Video Understanding and Editing
A Customizable Image-to-Video Model based on HunyuanVideo
Detect faces in an image
Blazeface is a lightweight model that detects faces in images
Implementation of Make-A-Video, new SOTA text to video generator
CoTracker is a model for tracking any point (pixel) on a video
A Strong and Easy-to-use Single View 3D Hand+Body Pose Estimator
UME is an in-app debug kits platform for Flutter