MiniMax H3 is an omni-modal generative system for creating synchronized video and native stereo audio from complex multimodal instructions. It can understand combinations of text, images, video, and audio as generation context. Outputs range from four to fifteen seconds at 24 frames per second, with multiple aspect ratios and resolutions up to 2K. H3-Base supports text-to-video, first-frame, last-frame, and first-and-last-frame generation. Ref2VA accepts multiple reference images, videos, and audio clips for more controlled results. Its architecture uses a unified multimodal sequence and a 33-billion-parameter H3-Omni-Transformer that jointly predicts video and audio latents. The repository also provides model components, inference code, prompt guidance, and specialized generation skills.

Features

  • Unified text, image, video, and audio understanding
  • Video generation with native stereo audio
  • Outputs up to 2K resolution
  • First-frame and last-frame conditioning
  • Multi-reference image, video, and audio input
  • 33-billion-parameter omni-modal Transformer

Project Samples

Project Activity

See All Activity >

Categories

AI Models

Follow MiniMax H3

MiniMax H3 Web Site

Other Useful Business Software
Build Agents and Models on One Platform Icon
Build Agents and Models on One Platform

Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
Try It Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of MiniMax H3!

Additional Project Details

Programming Language

Python

Related Categories

Python AI Models

Registered

2026-08-10