MiniMax H3

H3 Max

by MiniMax · Current; commercially hosted by fal

MiniMax H3 Max is fal's commercially hosted video-generation model based on MiniMax H3. It supports text, image, video, and audio-aware workflows, creates 5-to-15-second clips with synchronized audio, offers 480p to 1080p output, and charges per generated second.

Video generation Audio Reasoning Coding
MiniMax H3 Max is a commercially hosted video-generation model developed by fal from MiniMax H3. It is designed for fast text-to-video, image-to-video, and reference-guided production, with synchronized audio generated as part of the result. The model is aimed at short advertising clips, product visuals, cinematic scenes, social content, and rapid creative iteration rather than general-purpose language or coding tasks.
Outputs

What H3 Max can produce

Video generation Audio
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
10/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MiniMax H3
Model type Multimodal
Release date 2026-08-27
Status Current; commercially hosted by fal
Model notes

H3 Max is fal's post-trained variant of the open-weight MiniMax H3 model rather than a separately documented native MiniMax foundation-model release. It is optimized for prompt adherence, aesthetics, speed, and audio-visual generation. The model is available through fal-hosted endpoints including text-to-video, image-to-video, reference-to-video, camera controls, 3D-to-video, and video extension. Standard pricing after the September 14, 2026 promotional period is $0.05/sec at 480p, $0.08/sec at 768p, and $0.16/sec at 1080p. Five-to-fifteen-second clips are supported. Context length, maximum output tokens, knowledge cutoff, fine-tuning, caching, and batch API support have not been publicly specified for this video model.

Cost

Model pricing

Output $0.05 per second at 480p; $0.08 per second at 768p; $0.16 per second at 1080p
Model guide

MiniMax H3 Max: Fast Short-Form Video with Synchronized Audio

MiniMax H3 Max is a fal-hosted video-generation model post-trained from the open-weight MiniMax H3 model. It creates 5-to-15-second videos from text and multimodal references, supports up to 1080p output, and generates synchronized audio alongside the video. Its main advantage is fast production at a per-second price, while its main trade-offs are hosted-only access, short clip lengths, and limited suitability for language-model, coding, or self-hosted workloads.

What is MiniMax H3 Max?

MiniMax H3 Max is a video-generation model provided through fal and based on post-training of the open-weight MiniMax H3 model. In practical terms, it turns written prompts and selected reference media into short video clips. Depending on the endpoint, the input can include text, an image, video, or audio.

H3 Max is not simply the standard MiniMax H3 model exposed under another name. It is fal's commercially hosted production variant, optimized for prompt adherence, visual aesthetics, generation speed, and audio-visual output. The model is therefore best understood as a focused video-production endpoint rather than as a general-purpose multimodal assistant.

The model was released on August 27, 2026, and is accessed through fal-hosted endpoints rather than documented as a downloadable model for local inference.

Core capabilities and supported inputs

H3 Max supports several short-form video workflows. Text-to-video creates a scene from a written description, while image-to-video animates or transforms a supplied image. Reference-to-video can use additional media to guide the appearance or motion of the result. Separate fal workflows also expose camera controls, 3D-to-video generation, and video extension.

The supported input types vary by endpoint, but the documented model family supports:

  • Text prompts for describing the scene, action, style, camera behavior, and audio requirements.
  • Images for image-to-video animation and visual conditioning.
  • Video references for reference-guided generation and video extension workflows.
  • Audio references or audio-aware prompts in workflows that use sound as part of the conditioning or creative direction.

H3 Max can generate synchronized audio with the video, including environmental sound and dialogue-oriented audio when requested in the prompt. This distinguishes it from video systems that produce silent footage and require a separate audio-production step.

Output specifications and limits

H3 Max generates clips between 5 and 15 seconds long. The fal endpoint documents 480p, 768p, and 1080p output options, with 768p serving as the main balance between quality and performance. The model is therefore suited to short scenes and advertising or social-media segments, but it is not positioned as a long-form video-production system.

For text-to-video, documented aspect-ratio options include 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Image-to-video output generally follows the aspect ratio of the supplied image. This makes the model usable for both widescreen footage and vertical formats, although the final composition remains dependent on the source image or prompt.

SpecificationDocumented support
Video duration5 to 15 seconds
Output resolutions480p, 768p, and 1080p
Text-to-video aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, and 9:16
Output modalitiesVideo with synchronized audio
Context length and maximum output tokensNot applicable or not publicly specified for this video model

Traditional language-model specifications such as context length, maximum output tokens, and knowledge cutoff do not provide a useful way to describe H3 Max. The supplied documentation does not publish equivalent limits for this video endpoint.

Speed and pricing

fal states that H3 Max can generate a five-second video in under three seconds in typical configurations. This is a provider claim rather than a guaranteed result for every prompt, resolution, queue condition, or endpoint. Generation time can vary with the selected workflow and output settings.

Current standard pricing is charged per second of generated video:

ResolutionPrice per generated secondExample cost for 5 secondsExample cost for 15 seconds
480p$0.05$0.25$0.75
768p$0.08$0.40$1.20
1080p$0.16$0.80$2.40

These examples describe hosted fal inference and assume that the full requested duration is generated. They should not be confused with MiniMax consumer subscriptions, MiniMax Token Plan pricing, or pricing for the base MiniMax H3 model. Promotional launch rates that applied through September 14, 2026, are not the current standard rates described here.

The pricing creates a clear quality-versus-cost choice. 480p is the least expensive option for rapid drafts and large-volume experimentation. 768p is the documented middle tier and may be the most practical default for production iteration. 1080p costs twice as much per second as 768p, so it is better reserved for outputs where the additional resolution justifies the higher inference cost.

Synchronized audio and multimodal generation

H3 Max produces video and synchronized audio rather than text, images, or embeddings. Audio can include environmental effects and dialogue-oriented elements when the prompt calls for them. For a product demonstration, for example, a prompt might describe the product movement, camera framing, background setting, and the sound of the environment in one generation request.

The audio capability is useful for early creative reviews because the visual and sound tracks are produced together. It does not mean that the model replaces a full audio-post-production workflow. The supplied research does not specify detailed controls for mixing, editing, individual audio tracks, voice identity, or professional mastering.

Reasoning, coding, and tool support

H3 Max is a generative video model, not a reasoning-oriented language model. Its strength is interpreting creative instructions and producing a visual sequence, not solving multi-step analytical problems or returning carefully structured textual reasoning. The supplied model data assigns a low reasoning score, but that score is an editorial evaluation rather than a provider-published benchmark.

Coding is not a primary capability. The model data also assigns a low coding score, which should be treated as an editorial assessment rather than a formal provider rating. There is no documented general-purpose code execution or function-calling interface for H3 Max. The model is therefore not an appropriate substitute for a coding assistant or an agentic language model.

Streaming is listed in the model data, but the research does not describe the exact streaming behavior or whether it applies uniformly across every H3 Max workflow. Users should verify the behavior of the specific fal endpoint they intend to use.

Main strengths and limitations

Strengths

  • Fast iteration: fal claims that a five-second clip can be generated in under three seconds in typical configurations.
  • Native synchronized audio: video and audio can be generated together instead of requiring a separate silent-video workflow.
  • Multiple conditioning modes: text, images, video references, and audio-aware workflows support more than simple text-to-video generation.
  • Flexible formats: the documented aspect ratios include common landscape, square, portrait, and ultrawide formats.
  • Clear usage-based pricing: per-second billing makes the cost of a short test or production clip relatively easy to estimate.

Limitations

  • Short output duration: each generation is limited to 5-to-15-second clips, so longer sequences require multiple generations and editing.
  • Hosted access: H3 Max is presented as a fal-hosted production endpoint rather than a model for local or self-hosted inference.
  • No 2K option in the supplied specifications: the documented maximum is 1080p.
  • Separate workflows: text-to-video, image-to-video, reference-to-video, camera controls, 3D-to-video, and video extension are exposed through different fal endpoints.
  • Limited language-model applicability: context length, token limits, JSON output, general tool use, and fine-tuning are not documented as standard H3 Max features.
  • Output variability: fast generation does not guarantee that every prompt will produce consistent motion, composition, dialogue, or audio alignment.

Best use cases for H3 Max

H3 Max is a strong fit when the goal is to create short, visually directed clips quickly and review several variations. Suitable applications include:

  • Advertising concepts and branded social-media videos.
  • Product demonstrations and image-to-video product visualization.
  • Short cinematic scenes with environmental sound.
  • Storyboards and visual prototypes for games, film, or motion design.
  • Reference-guided creative production using supplied images or video.
  • Vertical clips for mobile-oriented publishing.

The model is particularly useful in workflows where speed matters more than maximum resolution and where synchronized sound is valuable during the first creative pass.

When to choose H3 Max

Choose H3 Max when you need short video generation through a hosted service, want to combine visual references with text instructions, and would benefit from audio generated alongside the footage. The 768p tier is a reasonable starting point for balancing output quality and cost, while 480p is better for inexpensive experiments and 1080p is intended for cases where additional resolution is important.

Another video model may be more appropriate when you need clips longer than 15 seconds in one generation, 2K output, extensive editing controls, or local deployment. The base MiniMax H3 model may be worth evaluating when self-hosting or access to the open-weight model is more important than fal's optimized hosted workflow. A language model is the better choice for coding, structured text generation, document reasoning, or general tool-using agents.

H3 Max should also be evaluated endpoint by endpoint. A workflow that supports image-to-video may not expose exactly the same controls, limits, or input handling as the text-to-video endpoint. Checking the selected fal endpoint is therefore important before estimating cost or building an automated production process.

Bottom line

MiniMax H3 Max is a focused solution for fast, short-form video generation with synchronized audio. Its strongest practical distinction is the combination of multimodal conditioning, several video-generation workflows, and per-second hosted pricing. It is less suitable for self-hosted deployment, long-form production, 2K-first projects, or general-purpose reasoning and coding. For creators and developers who need quick 5-to-15-second video variations, especially with sound and reference media, it offers a clearly defined speed and cost trade-off.


Answers to Frequently Asked Questions

How much does MiniMax H3 Max cost through fal?
Standard fal pricing is charged per generated second: $0.05 per second at 480p, $0.08 per second at 768p, and $0.16 per second at 1080p. A 5-second clip therefore costs $0.25, $0.40, or $0.80 respectively, while a 15-second clip costs $0.75, $1.20, or $2.40.
Does MiniMax H3 Max generate synchronized audio?
Yes. MiniMax H3 Max can generate video with synchronized audio, including environmental sounds and dialogue-oriented audio when requested. However, the available documentation does not specify advanced controls for mixing, individual audio tracks, voice identity, or professional mastering.
How long can videos generated by MiniMax H3 Max be?
MiniMax H3 Max generates videos between 5 and 15 seconds long. Longer sequences require multiple generations and editing.
What resolutions and aspect ratios does MiniMax H3 Max support?
The documented output resolutions are 480p, 768p, and 1080p. Text-to-video workflows support 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios, while image-to-video outputs generally follow the supplied image's aspect ratio.
What is MiniMax H3 Max?
MiniMax H3 Max is a fal-hosted video-generation model that turns text prompts and reference media such as images, video, or audio into short video clips with synchronized audio. It is a commercially hosted production variant based on post-training of the open-weight MiniMax H3 model.


Sources 10
Provider

About MiniMax