MiniMax H3 is a general-purpose, omni-modal generative system
Repo for SeedVR2 & SeedVR
Inference script for Oasis 500M
MiniMax H3 inference engine for Mac computers
Lets make video diffusion practical
Official Python inference and LoRA trainer package
Multimodal Diffusion with Representation Alignment
A Customizable Image-to-Video Model based on HunyuanVideo
Official repository for LTX-Video
Uncommon Objects in 3D dataset
Open-source multi-speaker long-form text-to-speech model
Inference code for scalable emulation of protein equilibrium ensembles
Code for running inference with the SAM 3D Body Model 3DB
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Video understanding codebase from FAIR for reproducing video models
Qwen2.5-VL is the multimodal large language model series
Large Multimodal Models for Video Understanding and Editing
OCR expert VLM powered by Hunyuan's native multimodal architecture
Tooling for the Common Objects In 3D dataset
Detect faces in an image
Blazeface is a lightweight model that detects faces in images
Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video
Multimodal 7B model for image, video, and text understanding tasks
Google’s flagship dense multimodal model for coding and reasoning