MiniMax H3 is a general-purpose, omni-modal generative system
Inference script for Oasis 500M
MiniMax H3 inference engine for Mac computers
Official Python inference and LoRA trainer package
A Customizable Image-to-Video Model based on HunyuanVideo
Multimodal Diffusion with Representation Alignment
Official repository for LTX-Video
Open-source multi-speaker long-form text-to-speech model
Code for running inference with the SAM 3D Body Model 3DB
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Video understanding codebase from FAIR for reproducing video models
Large Multimodal Models for Video Understanding and Editing
Detect faces in an image
Blazeface is a lightweight model that detects faces in images
Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video
Multimodal 7B model for image, video, and text understanding tasks
Google’s flagship dense multimodal model for coding and reasoning