RepoMiniMaxMiniMaxpublished Jul 30, 2026seen 4w

MiniMax-AI/MiniMax-H3

Python

Open original ↗

Captured source

source ↗
published Jul 30, 2026seen 4wcaptured 4whttp 200method plain

MiniMax-AI/MiniMax-H3

Language: Python

Stars: 298

Forks: 18

Open issues: 0

Created: 2026-07-30T08:41:09Z

Pushed: 2026-08-06T05:33:54Z

Default branch: main

Fork: no

Archived: no

README:

MiniMax H3

Prompt Writing Skill

Install the H3 prompt writing skill — one of nine skills bundled with this repository:

npx skills add https://github.com/MiniMax-AI/MiniMax-H3 --skill h3-prompt-writing

It ships with two prompt guides under skills/h3-prompt-writing/references/: base-en.txt for text/keyframe modes and ref-en.txt for full-reference (Ref2VA) mode. The remaining eight are style-specific video generation skills:

minimalist-product-ad-generator

3d-animation-short-generator

papercraft-stop-motion-explainer

brand-promo-video-generator

music-video-subtitle-generator

co-op-game-intro-generator

paper-collage-explainer-generator

handdrawn-live-video-generator

Online API

Use MiniMax\-H3 directly via API\.

Online App

Use MiniMax\-H3 directly via App\.

System Overview

MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions.

H3 supports the following input and output specifications:

| Category | Specification | |---|---| | Output duration | 4–15 seconds | | Output aspect ratio | Supports a wide range of aspect ratios, including but not limited to 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 | | Output resolution | Supports various resolution dimensions. The shorter side is set to 768 pixels by default. 2K \| generation can be achieved with H3-Regenerate-2K | | Output frame rate | 24 FPS | | Output audio | 32 kHz stereo | | Supported dialogue languages | Stable support for 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish. Additional languages are also supported to varying degrees |

Model Variants and Input Specifications

| Model Variant | Input Mode | Specifications | |---|---|---| | H3-Base-FL2VA | First-and-last-frame mode | Supports zero, one, or two input images.

  • No image input: Text-to-video mode
  • One image input: First-frame-to-video or last-frame-to-video generation
  • Two image inputs: First-and-last-frame-to-video generation |

| H3-Base-Ref2VA | Omni-reference mode | Supports multi-modal reference inputs:

  • Images: ≤ 9 images
  • Videos: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
  • Audio: ≤ 3 clips; audio must be accompanied by image or video input and cannot be used as the sole input; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
  • Mixed inputs: Maximum number of files across all input types is 12 |

![Image](assets/overview.png)

The complete H3 system consists of the following three modules:

  • H3-Context-IR: As inputs become increasingly complex, we build a dedicated system to deeply understand and refine the input multimodal instructions, then convert them into a form that H3 can readily understand—the Context Intermediate Representation—for generation. H3-Context-IR is critical to the quality of the final output, so we strongly recommend incorporating it into your generation pipeline or following the “Prompting Guidance” to build your own context-processing system.
  • H3-Base: Generates audio and video based on the H3-Context-IR output, producing results at 768p resolution.
  • H3-Regenerate-2K: Feeds the 768p result together with the original context back into H3 to regenerate the output at 2K resolution. This process leverages both H3’s powerful generative capabilities and the rich information contained in the original context, enabling it to produce high-resolution outputs with more accurate details and greater visual fidelity.

Model Architecture

H3\-Context\-IR

H3\-Context\-IR is a hosted preprocessing and orchestration system designed for free\-form multimodal inputs\.

It interprets the relationships among text, images, audio, and reference videos, as well as how these materials relate to the intended generation output\. Its internal workflow includes instruction parsing, cross\-modal association, temporal understanding, and complex logical reasoning\.

H3\-Context\-IR serializes its understanding of the context into a structured representation accepted by H3\-Base\. Without deviating from the user’s original intent, it may also supplement missing or underspecified semantic details where appropriate\.

Because H3\-Context\-IR relies on a multi\-stage workflow and multiple hosted models and services, it is not included in this open\-source release\. We provide an API that enables users to reproduce the behavior of the official workflow\. We also provide detailed tutorials, and developers can follow the Prompting Guidance to build their own preprocessing systems\.

For detailed usage instructions, see Recommended Workflow — Full 2K Workflow\.

Safety Guardrails

User\-submitted text, images and videos, as well as enhanced prompts, are subject to automated moderation\. Content suspected of being unlawful, pornographic, or infringing third\-party rights may be blocked\. We use industry\-standard filtering measures but cannot eliminate false positives or false negatives\. These guardrails do not affect the Licensee’s obligations under the MiniMax H3 Community License, especially those relating to lawful use and use restrictions\.

H3\-Base

![Image](assets/full-arch.png)

Architecture Overview

  • H3\-Base encodes different modalities using their corresponding encoders or VAEs and organizes the encoded representations into a unified packed multimodal sequence\. RoPE...

Excerpt shown — open the source for the full document.