MiniMax H3

MiniMax H3

by MiniMax · Current; open-weight release and hosted API available

MiniMax H3 is an open-weight multimodal video-generation system that accepts text, image, video, and audio references and produces short video with synchronized stereo audio. It supports 4–15-second clips, local 768p generation, hosted 2K output, first and last-frame control, reference-based editing, and separate Context-IR and regeneration services.

Video generation Music Reasoning Coding
MiniMax H3 is designed for creators and developers who need more than text-to-video generation. It can combine written instructions with visual and audio references, preserve subjects or motion across a clip, and generate video together with native stereo sound. The open-weight release supports local 768p workflows, while MiniMax's hosted API adds a 2K regeneration path.
Outputs

What MiniMax H3 can produce

Video generation Music
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MiniMax H3
Model type Multimodal
Release date 2026-07-31
Status Current; open-weight release and hosted API available
Knowledge cutoff notes

MiniMax has not published a direct knowledge-cutoff date for H3. The model is primarily a multimodal video-and-audio generation system, so a conventional language-model knowledge cutoff is not clearly applicable.

Model notes

MiniMax H3 is a video-and-audio generation system rather than a conventional text LLM. The open release provides H3-Base-FL2VA and H3-Base-Ref2VA checkpoints for local 768p generation. The full hosted 2K workflow additionally uses H3-Context-IR and H3-Regenerate-2K, which were not included in the initial open-weight release. H3-Omni-Transformer is a 33B dense transformer. H3 supports 4–15 second outputs, 24 FPS video, and 32 kHz stereo audio. The initial open release uses full attention; MiniMax has stated that sparse-attention inference will be released separately. Editorial scores reflect comparison with current video-generation systems, not text-model benchmarks.

Cost

Model pricing

Input $0.08 per second for 768p video output; $0.13 per second for 2K video output; H3-Context-IR: $0.90 per million input tokens
Output $0.08 per second for 768p video output; $0.13 per second for 2K video output; H3-Regenerate-2K: $0.05 per second of regenerated output; H3-Context-IR: $3.60 per million output tokens
Model guide

MiniMax H3: Open-Weight Video Generation with Native Stereo Audio

MiniMax H3 is an open-weight multimodal video-generation system from MiniMax. It accepts text, images, video, and audio references, and produces short videos with synchronized stereo audio. Developers can run the H3-Base checkpoints locally for 768p generation, while MiniMax's hosted workflow uses Context-IR and Regenerate-2K components to support output up to 2K resolution.

What is MiniMax H3?

MiniMax H3 is MiniMax's general-purpose multimodal video-generation model. Its main purpose is to create short commercial or creative video clips from combinations of text, images, video, and audio. Unlike a conventional text-to-video system that primarily turns a prompt into silent footage, H3 is designed to interpret relationships between several input modalities and generate video with synchronized stereo audio.

MiniMax released H3 as an open-weight model on July 31, 2026, followed by an open-source announcement on August 3, 2026. The model is available through the MiniMax Open Platform and through H3-Base checkpoints intended for local deployment. MiniMax positions it for advertising, branding, e-commerce, product design, UI/UX, gaming, and other commercial content workflows.

H3 is not a general-purpose chat model or a conventional text-generation endpoint. Its output is primarily video accompanied by audio, so it should be evaluated as a video-and-sound generation system rather than by the usual language-model measures such as text reasoning or token generation.

What can MiniMax H3 create?

H3 supports several video-generation and editing patterns. A user can provide a text description for text-to-video generation, specify a first or last frame, or provide both ends of a shot. It can also use reference images, video clips, and audio clips to guide the result.

  • Text-to-video: Generate a short video and synchronized audio from a written instruction.
  • First-frame and last-frame control: Use a specified beginning or ending image to guide the shot.
  • First-and-last-frame generation: Define both endpoints and ask H3 to create the transition between them.
  • Reference-to-video: Use images, video, and audio as references for subjects, movement, dialogue, music, or visual continuity.
  • Video-to-video workflows: Apply motion transfer and multimodal editing techniques to existing footage.

Generated clips can be between 4 and 15 seconds long. Supported aspect ratios include 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, covering cinematic, standard, square, portrait, and mobile-oriented formats. H3 produces video at 24 frames per second and native stereo audio at 32 kHz, according to the supplied model notes.

Supported inputs and outputs

H3 accepts text, images, video clips, and audio clips. The model's reference workflow is multimodal rather than merely accepting different file types independently: a prompt can explain how a subject should move, an image can establish appearance, a video can provide motion, and an audio clip can influence dialogue, sound, or music.

For the Ref2VA workflow, the supplied documentation supports up to nine images, up to three video clips, and up to three audio clips, subject to duration and combination restrictions. Audio cannot be used as the only reference in that workflow. These limits apply to the described reference workflow; the research does not provide a general context-window size or a universal maximum duration for every input type.

The direct model output is video with synchronized stereo audio. Local open-weight inference produces 768p output, described as having a default short side of 768 pixels. The hosted workflow can produce up to 2K resolution through a separate regeneration stage. H3 does not primarily provide text, still-image, embedding, or speech-synthesis output.

CapabilityVerified detail
Text inputSupported
Image inputSupported, including reference images
Video inputSupported, including reference and video-to-video workflows
Audio inputSupported as a reference, but not as the sole Ref2VA input
Video outputSupported, 4–15 seconds, 24 FPS
Audio outputNative synchronized stereo audio at 32 kHz
Local resolution768p through H3-Base
Hosted resolutionUp to 2K through the hosted regeneration workflow

How the H3 system is structured

The H3 release is made up of several components rather than one interchangeable checkpoint that performs every step. H3-Base generates 768p video and audio. Two task-specific checkpoint families are described in the open release: H3-Base-FL2VA handles text-to-audio-video and first/last-frame-to-audio-video generation, while H3-Base-Ref2VA handles multimodal reference-to-audio-video generation.

The hosted workflow adds H3-Context-IR and H3-Regenerate-2K. Context-IR is a preprocessing and orchestration service that converts complex multimodal instructions into a structured intermediate representation. In practical terms, it helps organize the relationships among the prompt and reference materials before generation. Regenerate-2K uses the lower-resolution result together with the original context to produce a higher-resolution result.

MiniMax describes H3-Omni-Transformer as a 33-billion-parameter dense transformer. The complete open-weight release is intended to support further development, including fine-tuning. However, the initial open release uses full attention, and MiniMax has described sparse-attention inference as a future update. Local deployment therefore offers control and independence, but it is not necessarily a lightweight setup.

Pricing and access

MiniMax's supplied API pricing lists H3 video generation at $0.08 per second for 768p output and $0.13 per second for 2K output. The 2K price applies to the hosted high-resolution workflow, not to a fully local 2K checkpoint.

Other hosted components are billed separately. H3-Context-IR costs $0.90 per million input tokens and $3.60 per million output tokens. The H3-Regenerate-2K stage costs $0.05 per second of regenerated output. Image and video reference materials may incur additional charges, while the supplied pricing information lists audio input as free.

For example, a 10-second 768p generation has a base video-generation charge of $0.80 before any applicable reference-material or Context-IR charges. A 10-second 2K generation has a listed video-generation charge of $1.30, and the regeneration stage may add another $0.50 if all 10 seconds are regenerated. These examples describe the listed per-second rates and do not account for other usage components.

The open-weight route avoids the hosted per-second generation charge but requires suitable local hardware and operational work. It also does not include the complete hosted 2K pipeline. Developers choosing local deployment should therefore distinguish between access to the H3-Base weights and access to MiniMax's managed Context-IR and regeneration services.

Reasoning, coding, and tool support

H3 is not intended to compete with text-first reasoning or coding models. It can interpret natural-language instructions for video creation, but the supplied research does not establish a conventional language-model reasoning benchmark, coding benchmark, context window, maximum text output, function-calling interface, tool-use system, web search capability, or JSON-output mode.

The model's technical strength is multimodal conditioning: it can use text and reference media to determine what should appear, how it should move, and what accompanying sound should be generated. That is different from reasoning through a long software task, generating reliable source code, or returning structured business data. For those jobs, a dedicated text or coding model would generally be more appropriate.

H3 does support further development and fine-tuning through its open-weight release, according to the supplied specifications. This is useful for teams that want to adapt a video-generation workflow, but it should not be confused with a built-in agent or function-calling capability.

Main strengths and limitations

Where H3 is strong

  • Unified multimodal conditioning: Text, image, video, and audio references can contribute to one generation workflow.
  • Native synchronized sound: H3 generates stereo audio with the video instead of requiring every project to add sound in a separate step.
  • Reference control: First and last frames, multiple images, video clips, and audio references support more directed shots than a prompt-only workflow.
  • Open-weight local inference: H3-Base enables developers to run 768p generation locally and potentially fine-tune the system.
  • Commercially useful formats: Short durations and several aspect ratios suit advertising, product demonstrations, social video, and branded creative work.

Where H3 is limited

  • Short clips: The documented output duration is 4–15 seconds, so long-form scenes require multiple generations and editing.
  • Separate 2K pipeline: Local H3-Base generation is 768p; the hosted 2K path requires Context-IR and Regenerate-2K services.
  • Hardware demands: The 33-billion-parameter dense transformer and initial full-attention implementation may make local deployment demanding.
  • No text-model role: H3 is not a replacement for a chat, coding, embedding, speech, or general reasoning model.
  • Input-combination restrictions: Reference counts and duration limits apply, and audio cannot be the sole Ref2VA reference.
  • Potentially higher workflow complexity: The open checkpoints and hosted components have different roles, pricing, and deployment requirements.

The strengths and limitations above combine verified specifications with practical evaluation. The specification details come from MiniMax's supplied materials; statements about suitability for particular production workflows are editorial judgments based on those capabilities.

When to choose MiniMax H3

Choose H3 when the project needs short video with synchronized sound and benefits from several kinds of reference input. Suitable examples include a product demonstration guided by product images, a branded vertical clip based on a reference video and soundtrack, a cinematic transition defined by first and last frames, or a character shot that must preserve visual details while following an audio cue.

H3 is especially attractive when local 768p generation, open weights, or future fine-tuning matter. A development team can keep more control over the local generation workflow instead of sending every generation request to a hosted service. The trade-off is the need to manage substantial compute and accept that the hosted 2K components are not part of the initial open release.

Choose a different type of system when the priority is long-form video in one pass, text-heavy reasoning, software development, embeddings, speech synthesis, or a fully local 2K workflow. A simpler text-to-video tool may be preferable for prompt-only experiments where audio and multiple references are unnecessary. A dedicated editing or compositing pipeline may also be more efficient when precise frame-level control matters more than native generative audio.

Compared with other options in the current video-generation market, H3's distinguishing trade-off is breadth of conditioning and integrated stereo sound rather than a documented advantage in every quality or speed category. The supplied research does not provide benchmark results or a direct speed comparison with competing systems. Its practical value is therefore strongest for users who will actually use multimodal references, native audio, open weights, or the hosted 768p-to-2K workflow.

Bottom line

MiniMax H3 is a specialized video-and-audio generation system for short, reference-controlled clips. It combines text, image, video, and audio understanding with synchronized stereo output, offers local 768p open-weight inference, and provides hosted generation up to 2K through additional MiniMax services. Its main decision points are the need for native sound and multimodal references versus the cost, hardware, short-duration limits, and complexity of its multi-component deployment model.


Answers to Frequently Asked Questions

How much does MiniMax H3 video generation cost?
The listed hosted price is $0.08 per second for 768p video and $0.13 per second for 2K video. A 10-second 768p generation therefore costs $0.80 before additional charges, while a 10-second 2K generation costs $1.30, with the hosted Regenerate-2K stage potentially adding another $0.50.
Can MiniMax H3 be run locally?
Yes. MiniMax released H3-Base open-weight checkpoints for local deployment and potential fine-tuning. Local inference produces 768p video, but the 33-billion-parameter dense transformer and full-attention implementation may require substantial hardware. The hosted 2K generation pipeline is separate from the local open-weight release.
Does MiniMax H3 generate audio with video?
Yes. MiniMax H3 generates synchronized native stereo audio at 32 kHz along with the video. Audio clips can also be used as references, although audio cannot be the only reference input in the documented Ref2VA workflow.
What types of video can MiniMax H3 generate?
H3 supports text-to-video, first-frame and last-frame control, first-and-last-frame transitions, reference-to-video generation, and video-to-video workflows. Generated clips are documented at 4–15 seconds, 24 frames per second, with support for aspect ratios including 16:9, 9:16, 1:1, 4:3, 3:4, and 21:9.
What is MiniMax H3?
MiniMax H3 is a multimodal video-generation model that creates short video clips with synchronized stereo audio from text, images, video, and audio references. It is designed for creative and commercial workflows such as advertising, branding, e-commerce, gaming, and product design.


Sources 8
Provider

About MiniMax