What is MiniMax H3?
MiniMax H3 is MiniMax's general-purpose multimodal video-generation model. Its main purpose is to create short commercial or creative video clips from combinations of text, images, video, and audio. Unlike a conventional text-to-video system that primarily turns a prompt into silent footage, H3 is designed to interpret relationships between several input modalities and generate video with synchronized stereo audio.
MiniMax released H3 as an open-weight model on July 31, 2026, followed by an open-source announcement on August 3, 2026. The model is available through the MiniMax Open Platform and through H3-Base checkpoints intended for local deployment. MiniMax positions it for advertising, branding, e-commerce, product design, UI/UX, gaming, and other commercial content workflows.
H3 is not a general-purpose chat model or a conventional text-generation endpoint. Its output is primarily video accompanied by audio, so it should be evaluated as a video-and-sound generation system rather than by the usual language-model measures such as text reasoning or token generation.
What can MiniMax H3 create?
H3 supports several video-generation and editing patterns. A user can provide a text description for text-to-video generation, specify a first or last frame, or provide both ends of a shot. It can also use reference images, video clips, and audio clips to guide the result.
- Text-to-video: Generate a short video and synchronized audio from a written instruction.
- First-frame and last-frame control: Use a specified beginning or ending image to guide the shot.
- First-and-last-frame generation: Define both endpoints and ask H3 to create the transition between them.
- Reference-to-video: Use images, video, and audio as references for subjects, movement, dialogue, music, or visual continuity.
- Video-to-video workflows: Apply motion transfer and multimodal editing techniques to existing footage.
Generated clips can be between 4 and 15 seconds long. Supported aspect ratios include 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, covering cinematic, standard, square, portrait, and mobile-oriented formats. H3 produces video at 24 frames per second and native stereo audio at 32 kHz, according to the supplied model notes.
Supported inputs and outputs
H3 accepts text, images, video clips, and audio clips. The model's reference workflow is multimodal rather than merely accepting different file types independently: a prompt can explain how a subject should move, an image can establish appearance, a video can provide motion, and an audio clip can influence dialogue, sound, or music.
For the Ref2VA workflow, the supplied documentation supports up to nine images, up to three video clips, and up to three audio clips, subject to duration and combination restrictions. Audio cannot be used as the only reference in that workflow. These limits apply to the described reference workflow; the research does not provide a general context-window size or a universal maximum duration for every input type.
The direct model output is video with synchronized stereo audio. Local open-weight inference produces 768p output, described as having a default short side of 768 pixels. The hosted workflow can produce up to 2K resolution through a separate regeneration stage. H3 does not primarily provide text, still-image, embedding, or speech-synthesis output.
| Capability | Verified detail |
|---|---|
| Text input | Supported |
| Image input | Supported, including reference images |
| Video input | Supported, including reference and video-to-video workflows |
| Audio input | Supported as a reference, but not as the sole Ref2VA input |
| Video output | Supported, 4–15 seconds, 24 FPS |
| Audio output | Native synchronized stereo audio at 32 kHz |
| Local resolution | 768p through H3-Base |
| Hosted resolution | Up to 2K through the hosted regeneration workflow |
How the H3 system is structured
The H3 release is made up of several components rather than one interchangeable checkpoint that performs every step. H3-Base generates 768p video and audio. Two task-specific checkpoint families are described in the open release: H3-Base-FL2VA handles text-to-audio-video and first/last-frame-to-audio-video generation, while H3-Base-Ref2VA handles multimodal reference-to-audio-video generation.
The hosted workflow adds H3-Context-IR and H3-Regenerate-2K. Context-IR is a preprocessing and orchestration service that converts complex multimodal instructions into a structured intermediate representation. In practical terms, it helps organize the relationships among the prompt and reference materials before generation. Regenerate-2K uses the lower-resolution result together with the original context to produce a higher-resolution result.
MiniMax describes H3-Omni-Transformer as a 33-billion-parameter dense transformer. The complete open-weight release is intended to support further development, including fine-tuning. However, the initial open release uses full attention, and MiniMax has described sparse-attention inference as a future update. Local deployment therefore offers control and independence, but it is not necessarily a lightweight setup.
Pricing and access
MiniMax's supplied API pricing lists H3 video generation at $0.08 per second for 768p output and $0.13 per second for 2K output. The 2K price applies to the hosted high-resolution workflow, not to a fully local 2K checkpoint.
Other hosted components are billed separately. H3-Context-IR costs $0.90 per million input tokens and $3.60 per million output tokens. The H3-Regenerate-2K stage costs $0.05 per second of regenerated output. Image and video reference materials may incur additional charges, while the supplied pricing information lists audio input as free.
For example, a 10-second 768p generation has a base video-generation charge of $0.80 before any applicable reference-material or Context-IR charges. A 10-second 2K generation has a listed video-generation charge of $1.30, and the regeneration stage may add another $0.50 if all 10 seconds are regenerated. These examples describe the listed per-second rates and do not account for other usage components.
The open-weight route avoids the hosted per-second generation charge but requires suitable local hardware and operational work. It also does not include the complete hosted 2K pipeline. Developers choosing local deployment should therefore distinguish between access to the H3-Base weights and access to MiniMax's managed Context-IR and regeneration services.
Reasoning, coding, and tool support
H3 is not intended to compete with text-first reasoning or coding models. It can interpret natural-language instructions for video creation, but the supplied research does not establish a conventional language-model reasoning benchmark, coding benchmark, context window, maximum text output, function-calling interface, tool-use system, web search capability, or JSON-output mode.
The model's technical strength is multimodal conditioning: it can use text and reference media to determine what should appear, how it should move, and what accompanying sound should be generated. That is different from reasoning through a long software task, generating reliable source code, or returning structured business data. For those jobs, a dedicated text or coding model would generally be more appropriate.
H3 does support further development and fine-tuning through its open-weight release, according to the supplied specifications. This is useful for teams that want to adapt a video-generation workflow, but it should not be confused with a built-in agent or function-calling capability.
Main strengths and limitations
Where H3 is strong
- Unified multimodal conditioning: Text, image, video, and audio references can contribute to one generation workflow.
- Native synchronized sound: H3 generates stereo audio with the video instead of requiring every project to add sound in a separate step.
- Reference control: First and last frames, multiple images, video clips, and audio references support more directed shots than a prompt-only workflow.
- Open-weight local inference: H3-Base enables developers to run 768p generation locally and potentially fine-tune the system.
- Commercially useful formats: Short durations and several aspect ratios suit advertising, product demonstrations, social video, and branded creative work.
Where H3 is limited
- Short clips: The documented output duration is 4–15 seconds, so long-form scenes require multiple generations and editing.
- Separate 2K pipeline: Local H3-Base generation is 768p; the hosted 2K path requires Context-IR and Regenerate-2K services.
- Hardware demands: The 33-billion-parameter dense transformer and initial full-attention implementation may make local deployment demanding.
- No text-model role: H3 is not a replacement for a chat, coding, embedding, speech, or general reasoning model.
- Input-combination restrictions: Reference counts and duration limits apply, and audio cannot be the sole Ref2VA reference.
- Potentially higher workflow complexity: The open checkpoints and hosted components have different roles, pricing, and deployment requirements.
The strengths and limitations above combine verified specifications with practical evaluation. The specification details come from MiniMax's supplied materials; statements about suitability for particular production workflows are editorial judgments based on those capabilities.
When to choose MiniMax H3
Choose H3 when the project needs short video with synchronized sound and benefits from several kinds of reference input. Suitable examples include a product demonstration guided by product images, a branded vertical clip based on a reference video and soundtrack, a cinematic transition defined by first and last frames, or a character shot that must preserve visual details while following an audio cue.
H3 is especially attractive when local 768p generation, open weights, or future fine-tuning matter. A development team can keep more control over the local generation workflow instead of sending every generation request to a hosted service. The trade-off is the need to manage substantial compute and accept that the hosted 2K components are not part of the initial open release.
Choose a different type of system when the priority is long-form video in one pass, text-heavy reasoning, software development, embeddings, speech synthesis, or a fully local 2K workflow. A simpler text-to-video tool may be preferable for prompt-only experiments where audio and multiple references are unnecessary. A dedicated editing or compositing pipeline may also be more efficient when precise frame-level control matters more than native generative audio.
Compared with other options in the current video-generation market, H3's distinguishing trade-off is breadth of conditioning and integrated stereo sound rather than a documented advantage in every quality or speed category. The supplied research does not provide benchmark results or a direct speed comparison with competing systems. Its practical value is therefore strongest for users who will actually use multimodal references, native audio, open weights, or the hosted 768p-to-2K workflow.
Bottom line
MiniMax H3 is a specialized video-and-audio generation system for short, reference-controlled clips. It combines text, image, video, and audio understanding with synchronized stereo output, offers local 768p open-weight inference, and provides hosted generation up to 2K through additional MiniMax services. Its main decision points are the need for native sound and multimodal references versus the cost, hardware, short-duration limits, and complexity of its multi-component deployment model.

