MiniMax H3 is a general-purpose, omni-modal generative system
Inference script for Oasis 500M
MiniMax H3 inference engine for Mac computers
Official Python inference and LoRA trainer package
A Customizable Image-to-Video Model based on HunyuanVideo
Lets make video diffusion practical
Multimodal Diffusion with Representation Alignment
Official repository for LTX-Video
Open-source multi-speaker long-form text-to-speech model
Code for running inference with the SAM 3D Body Model 3DB
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Qwen2.5-VL is the multimodal large language model series
Video understanding codebase from FAIR for reproducing video models
Large Multimodal Models for Video Understanding and Editing
Tooling for the Common Objects In 3D dataset
Detect faces in an image
Blazeface is a lightweight model that detects faces in images
Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video
Multimodal 7B model for image, video, and text understanding tasks
Google’s flagship dense multimodal model for coding and reasoning