Hailuo

MiniMax Hailuo 2.3

by MiniMax · Legacy or superseded video model; still referenced in MiniMax consumer and platform offerings, while MiniMax H3 is the current primary video model in developer documentation.

MiniMax Hailuo 2.3 is a short-form video generation model for text-to-video and image-to-video creation. It supports 768P and 1080P output, realistic motion, expressive characters, cinematic effects, and stylized visual generation, with six- and ten-second clip options. The model is now a legacy or superseded choice as MiniMax H3 becomes more prominent in current developer documentation.

Video generation Reasoning Coding
MiniMax Hailuo 2.3 is a closed video-generation model for creating short cinematic clips from text prompts or reference images. It is aimed at advertising, animation, visual effects, social-media content, and rapid creative iteration. The standard model supports both text-to-video and image-to-video generation, with six- and ten-second options at 768P and six-second output at 1080P. A separate Hailuo 2.3 Fast variant prioritizes speed and cost, but it is not the same model and has narrower input support.
Outputs

What MiniMax Hailuo 2.3 can produce

Video generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Hailuo
Model type Video Generation
Release date 2025-10-28
Status Legacy or superseded video model; still referenced in MiniMax consumer and platform offerings, while MiniMax H3 is the current primary video model in developer documentation.
Knowledge cutoff notes

A knowledge cutoff is not applicable or publicly documented for this generative video model.

Model notes

The canonical model identifier is generally written as MiniMax-Hailuo-2.3. The model supports text-to-video and image-to-video generation, with 768P and 1080P output. Six-second generation is available at both resolutions; ten-second generation is limited to 768P. MiniMax launched a separate Hailuo 2.3 Fast variant with a speed and cost optimization focus, but it is not the same exact model and is not represented by this record. Pricing is generation-based rather than token-based and may vary by resource package, promotion, region, or platform. MiniMax H3 is the newer primary video model in current developer documentation.

Cost

Model pricing

Output $0.28 per 768P/6s clip; $0.56 per 768P/10s clip; $0.49 per 1080P/6s clip
Model guide

MiniMax Hailuo 2.3 for Realistic Motion and Stylized AI Video

MiniMax Hailuo 2.3 is a short-form video generation model released on October 28, 2025. It supports text-to-video and image-to-video workflows, producing six- or ten-second clips at 768P and six-second clips at 1080P. Its main focus is realistic human movement, facial expression, camera motion, visual effects, and stylized scenes. Hailuo 2.3 is now a legacy or superseded model in MiniMax's lineup, with H3 appearing as the newer primary video model in current developer documentation.

What is MiniMax Hailuo 2.3?

MiniMax Hailuo 2.3 is a short-form generative video model from MiniMax. Released on October 28, 2025, it is designed to turn written descriptions or still images into short video clips. Unlike a general-purpose language model, it does not primarily answer questions, write code, or produce text. Its output is video.

The model is an updated member of the Hailuo video family and builds on Hailuo 02. MiniMax describes the release as an improvement in physical motion, facial micro-expressions, prompt adherence, camera movement, visual consistency, and stylized rendering. These are provider claims about the model's intended improvements; the supplied research does not include an independent benchmark that quantifies them.

Hailuo 2.3 is most relevant to creators who need a short visual sequence rather than a long-form production. A user can describe a scene such as a person walking through a rain-soaked street, or provide an image of a character and ask the system to animate it. The result is a brief clip suitable for concept development, advertising drafts, social content, visual effects experiments, and other short-video workflows.

Inputs, outputs, and supported formats

The standard Hailuo 2.3 model supports two primary workflows:

  • Text-to-video: Generate a clip from a written description of the subject, setting, action, camera direction, or visual style.
  • Image-to-video: Supply a reference image and animate it into a moving scene.

Its documented output is video only. The available research does not confirm native audio, music, speech, image, text, embedding, or structured-data output for the exact model. Audio may be relevant to a broader video product workflow, but it should not be assumed to be generated by Hailuo 2.3 itself.

MiniMax positions the model for realistic human motion, complex body actions, dynamic camera movement, visual effects, and expressive faces. It also advertises support for a range of visual styles, including anime, illustration, ink-wash painting, game computer graphics, and cinematic imagery. These style descriptions indicate the intended creative range, not a guarantee that every prompt will produce consistent results.

Technical specifications and clip limits

SpecificationHailuo 2.3
ProviderMiniMax
Release dateOctober 28, 2025
Model typeVideo generation
Input typesText and image
Output typeVideo
Supported resolutions768P and 1080P
768P duration options6 seconds or 10 seconds
1080P duration option6 seconds
Audio generationNot documented for the exact model
Context length and token limitNot applicable or not publicly documented

The duration and resolution limits have practical consequences. A ten-second clip is available at 768P, while 1080P output is limited to six seconds. Users who need higher resolution must therefore work with a shorter shot. Conversely, users who need a longer short-form sequence must accept the lower resolution option or assemble multiple clips in an editing workflow.

The model does not have a conventional language-model context window or maximum output-token limit. The available documentation also does not specify a maximum prompt length, image dimensions, supported image formats, or a guaranteed number of simultaneous generations. Those details should be checked in the active MiniMax interface or API documentation before building a production workflow.

Where Hailuo 2.3 is strongest

Hailuo 2.3 is most differentiated by the type of motion and visual direction it is intended to handle. MiniMax highlights realistic physical movement, facial micro-expressions, complex body actions, camera movement, and visual effects. In practical terms, this makes the model a candidate for shots where a static image is not enough and the movement itself is central to the scene.

Potentially suitable examples include:

  • A product concept with a controlled camera move around an object.
  • A short advertisement showing a character interacting with a setting.
  • An image brought to life with subtle facial or body movement.
  • An action or dance concept requiring visible body motion.
  • A stylized animation test in an anime, illustrated, game-CG, or ink-wash style.
  • A cinematic establishing shot or visual-effects concept for early production planning.

Image-to-video support is particularly useful when the creator already has a character design, storyboard frame, product image, or concept illustration. It provides a starting visual reference rather than requiring the model to invent every detail from text alone. However, the supplied research does not establish a formal guarantee of character identity preservation or frame-to-frame consistency.

Pricing and availability

MiniMax launched Hailuo 2.3 through the Hailuo AI website, a mobile application, and the MiniMax Open Platform API. Reported pay-as-you-go API rates for the standard model are approximately $0.28 for a 768P six-second clip, $0.56 for a 768P ten-second clip, and $0.49 for a 1080P six-second clip.

GenerationReported price
768P, 6 secondsApproximately $0.28 per clip
768P, 10 secondsApproximately $0.56 per clip
1080P, 6 secondsApproximately $0.49 per clip

These are generation-based prices rather than token prices. They are reported API rates, not a guaranteed universal price for every account or access method. The final cost may vary with resource packages, promotions, region, account type, or later changes to MiniMax's platform. Consumer access may also use product-specific credits, quotas, queues, or other restrictions.

Before committing to an integration, users should confirm that Hailuo 2.3 is still exposed through the selected endpoint and check the current price for the exact resolution and duration. By September 2026, MiniMax's newer H3 video model had become the primary model shown in current developer documentation. Hailuo 2.3 therefore remains relevant for existing workflows and product access, but it should be treated as a legacy or superseded option rather than automatically assumed to be MiniMax's newest video model.

Hailuo 2.3 compared with the Fast variant and newer models

MiniMax also released Hailuo 2.3 Fast. The Fast version is a separate model focused on quicker and lower-cost generation, particularly for batch work. The available research says that it has narrower input support and is primarily positioned for image-to-video workflows. It should not be treated as an interchangeable name for the standard Hailuo 2.3 model.

The standard model is the better fit when text-to-video generation is important or when the broader advertised capability set matters more than the shortest turnaround. Hailuo 2.3 Fast may be more appropriate when a team has many reference images to animate and values throughput or cost over the full input range of the standard model. Exact Fast pricing and technical limits are not supplied here, so those details require separate verification.

MiniMax H3 is the newer primary video model in current developer documentation. That does not prove that H3 is better for every individual shot, but it does mean that users starting a new API integration should compare Hailuo 2.3 with H3 rather than selecting Hailuo 2.3 solely because it was previously prominent. Existing Hailuo workflows may still favor 2.3 when compatibility, an established creative style, or continued product availability is more important than adopting the newest documented model.

Reasoning, coding, and tool support

Hailuo 2.3 is a specialized video generator, not a reasoning or coding model. It has no published reasoning capability intended for multi-step text analysis, and it is not designed to generate or execute software. The model record assigns low editorial scores for reasoning and coding, but these are database evaluations rather than provider-published benchmark results.

No tool or function-calling capability is documented for the exact model. Streaming, fine-tuning, prompt caching, batch API support, and structured outputs are also not confirmed in the available research. The Open Platform may provide surrounding API controls, but platform-level features should not be attributed to Hailuo 2.3 unless the model-specific documentation confirms them.

Strengths and limitations

Strengths

  • Supports both text-to-video and image-to-video workflows.
  • Offers 768P and 1080P output choices.
  • Can produce six-second clips at both listed resolutions and ten-second clips at 768P.
  • Targets realistic human movement, facial expression, camera motion, and visual effects.
  • Supports creative styles such as anime, illustration, ink wash, game CG, and cinematic imagery.
  • Has reported per-clip API pricing that is easy to estimate for short-form experimentation.

Limitations

  • Output is restricted to short clips rather than long-form scenes.
  • Ten-second generation is limited to 768P, while 1080P is limited to six seconds.
  • Native audio or music generation is not documented for the exact model.
  • There is no documented text, image, or audio output mode.
  • Prompt limits, image constraints, concurrency, and several API controls are not publicly established in the supplied research.
  • Availability and pricing may change as MiniMax moves users toward newer models such as H3.

When should you choose Hailuo 2.3?

Choose Hailuo 2.3 when you need short video generation with both text and image inputs, especially when realistic motion, expressive characters, camera direction, or stylized rendering is more important than long duration. It is a sensible option for creative exploration, advertising concepts, image animation, social clips, and pre-production visuals where several short variations may be generated and edited together.

Consider another option when you need long-form or continuous video, native sound, documented function calling, fine-tuning, structured outputs, or a clearly supported current API model. Hailuo 2.3 Fast may be preferable for image-to-video batches where speed and cost are the main priorities. MiniMax H3 deserves comparison for new integrations because it is the newer primary video model in current developer documentation.

Overall, Hailuo 2.3 is best understood as a focused short-video model with a useful balance of image and text input support, cinematic motion targets, and predictable per-clip pricing. Its main trade-off is that the model is no longer the clearest forward-looking choice in MiniMax's catalog, and its short output limits make it a component for edited video workflows rather than a complete long-form production system.


Answers to Frequently Asked Questions

When should you choose Hailuo 2.3 instead of Hailuo 2.3 Fast or MiniMax H3?
Choose standard Hailuo 2.3 when you need both text-to-video and image-to-video generation with realistic motion or stylized rendering. Hailuo 2.3 Fast may suit image-to-video batch workflows that prioritize speed and cost, while MiniMax H3 should be compared for new integrations because it is the newer primary video model in current developer documentation.
How much does MiniMax Hailuo 2.3 cost?
Reported API prices are approximately $0.28 for a 768P six-second clip, $0.56 for a 768P ten-second clip, and $0.49 for a 1080P six-second clip. Actual prices may vary by account, region, promotions, resource packages, or later platform changes.
What resolutions and video durations are available in Hailuo 2.3?
Hailuo 2.3 supports 768P and 1080P video. At 768P, users can generate six- or ten-second clips; at 1080P, the documented duration is six seconds.
What is MiniMax Hailuo 2.3?
MiniMax Hailuo 2.3 is a short-form generative video model that creates video clips from text descriptions or still images. It is designed for realistic motion, facial expressions, camera movement, visual effects, and stylized video generation.
What inputs and outputs does Hailuo 2.3 support?
Hailuo 2.3 supports text-to-video and image-to-video workflows. Its documented output is video, with no confirmed native audio, music, image, text, embedding, or structured-data output for the exact model.


Sources 5
Provider

About MiniMax