What is MiniMax Hailuo 2.3?
MiniMax Hailuo 2.3 is a short-form generative video model from MiniMax. Released on October 28, 2025, it is designed to turn written descriptions or still images into short video clips. Unlike a general-purpose language model, it does not primarily answer questions, write code, or produce text. Its output is video.
The model is an updated member of the Hailuo video family and builds on Hailuo 02. MiniMax describes the release as an improvement in physical motion, facial micro-expressions, prompt adherence, camera movement, visual consistency, and stylized rendering. These are provider claims about the model's intended improvements; the supplied research does not include an independent benchmark that quantifies them.
Hailuo 2.3 is most relevant to creators who need a short visual sequence rather than a long-form production. A user can describe a scene such as a person walking through a rain-soaked street, or provide an image of a character and ask the system to animate it. The result is a brief clip suitable for concept development, advertising drafts, social content, visual effects experiments, and other short-video workflows.
Inputs, outputs, and supported formats
The standard Hailuo 2.3 model supports two primary workflows:
- Text-to-video: Generate a clip from a written description of the subject, setting, action, camera direction, or visual style.
- Image-to-video: Supply a reference image and animate it into a moving scene.
Its documented output is video only. The available research does not confirm native audio, music, speech, image, text, embedding, or structured-data output for the exact model. Audio may be relevant to a broader video product workflow, but it should not be assumed to be generated by Hailuo 2.3 itself.
MiniMax positions the model for realistic human motion, complex body actions, dynamic camera movement, visual effects, and expressive faces. It also advertises support for a range of visual styles, including anime, illustration, ink-wash painting, game computer graphics, and cinematic imagery. These style descriptions indicate the intended creative range, not a guarantee that every prompt will produce consistent results.
Technical specifications and clip limits
| Specification | Hailuo 2.3 |
|---|---|
| Provider | MiniMax |
| Release date | October 28, 2025 |
| Model type | Video generation |
| Input types | Text and image |
| Output type | Video |
| Supported resolutions | 768P and 1080P |
| 768P duration options | 6 seconds or 10 seconds |
| 1080P duration option | 6 seconds |
| Audio generation | Not documented for the exact model |
| Context length and token limit | Not applicable or not publicly documented |
The duration and resolution limits have practical consequences. A ten-second clip is available at 768P, while 1080P output is limited to six seconds. Users who need higher resolution must therefore work with a shorter shot. Conversely, users who need a longer short-form sequence must accept the lower resolution option or assemble multiple clips in an editing workflow.
The model does not have a conventional language-model context window or maximum output-token limit. The available documentation also does not specify a maximum prompt length, image dimensions, supported image formats, or a guaranteed number of simultaneous generations. Those details should be checked in the active MiniMax interface or API documentation before building a production workflow.
Where Hailuo 2.3 is strongest
Hailuo 2.3 is most differentiated by the type of motion and visual direction it is intended to handle. MiniMax highlights realistic physical movement, facial micro-expressions, complex body actions, camera movement, and visual effects. In practical terms, this makes the model a candidate for shots where a static image is not enough and the movement itself is central to the scene.
Potentially suitable examples include:
- A product concept with a controlled camera move around an object.
- A short advertisement showing a character interacting with a setting.
- An image brought to life with subtle facial or body movement.
- An action or dance concept requiring visible body motion.
- A stylized animation test in an anime, illustrated, game-CG, or ink-wash style.
- A cinematic establishing shot or visual-effects concept for early production planning.
Image-to-video support is particularly useful when the creator already has a character design, storyboard frame, product image, or concept illustration. It provides a starting visual reference rather than requiring the model to invent every detail from text alone. However, the supplied research does not establish a formal guarantee of character identity preservation or frame-to-frame consistency.
Pricing and availability
MiniMax launched Hailuo 2.3 through the Hailuo AI website, a mobile application, and the MiniMax Open Platform API. Reported pay-as-you-go API rates for the standard model are approximately $0.28 for a 768P six-second clip, $0.56 for a 768P ten-second clip, and $0.49 for a 1080P six-second clip.
| Generation | Reported price |
|---|---|
| 768P, 6 seconds | Approximately $0.28 per clip |
| 768P, 10 seconds | Approximately $0.56 per clip |
| 1080P, 6 seconds | Approximately $0.49 per clip |
These are generation-based prices rather than token prices. They are reported API rates, not a guaranteed universal price for every account or access method. The final cost may vary with resource packages, promotions, region, account type, or later changes to MiniMax's platform. Consumer access may also use product-specific credits, quotas, queues, or other restrictions.
Before committing to an integration, users should confirm that Hailuo 2.3 is still exposed through the selected endpoint and check the current price for the exact resolution and duration. By September 2026, MiniMax's newer H3 video model had become the primary model shown in current developer documentation. Hailuo 2.3 therefore remains relevant for existing workflows and product access, but it should be treated as a legacy or superseded option rather than automatically assumed to be MiniMax's newest video model.
Hailuo 2.3 compared with the Fast variant and newer models
MiniMax also released Hailuo 2.3 Fast. The Fast version is a separate model focused on quicker and lower-cost generation, particularly for batch work. The available research says that it has narrower input support and is primarily positioned for image-to-video workflows. It should not be treated as an interchangeable name for the standard Hailuo 2.3 model.
The standard model is the better fit when text-to-video generation is important or when the broader advertised capability set matters more than the shortest turnaround. Hailuo 2.3 Fast may be more appropriate when a team has many reference images to animate and values throughput or cost over the full input range of the standard model. Exact Fast pricing and technical limits are not supplied here, so those details require separate verification.
MiniMax H3 is the newer primary video model in current developer documentation. That does not prove that H3 is better for every individual shot, but it does mean that users starting a new API integration should compare Hailuo 2.3 with H3 rather than selecting Hailuo 2.3 solely because it was previously prominent. Existing Hailuo workflows may still favor 2.3 when compatibility, an established creative style, or continued product availability is more important than adopting the newest documented model.
Reasoning, coding, and tool support
Hailuo 2.3 is a specialized video generator, not a reasoning or coding model. It has no published reasoning capability intended for multi-step text analysis, and it is not designed to generate or execute software. The model record assigns low editorial scores for reasoning and coding, but these are database evaluations rather than provider-published benchmark results.
No tool or function-calling capability is documented for the exact model. Streaming, fine-tuning, prompt caching, batch API support, and structured outputs are also not confirmed in the available research. The Open Platform may provide surrounding API controls, but platform-level features should not be attributed to Hailuo 2.3 unless the model-specific documentation confirms them.
Strengths and limitations
Strengths
- Supports both text-to-video and image-to-video workflows.
- Offers 768P and 1080P output choices.
- Can produce six-second clips at both listed resolutions and ten-second clips at 768P.
- Targets realistic human movement, facial expression, camera motion, and visual effects.
- Supports creative styles such as anime, illustration, ink wash, game CG, and cinematic imagery.
- Has reported per-clip API pricing that is easy to estimate for short-form experimentation.
Limitations
- Output is restricted to short clips rather than long-form scenes.
- Ten-second generation is limited to 768P, while 1080P is limited to six seconds.
- Native audio or music generation is not documented for the exact model.
- There is no documented text, image, or audio output mode.
- Prompt limits, image constraints, concurrency, and several API controls are not publicly established in the supplied research.
- Availability and pricing may change as MiniMax moves users toward newer models such as H3.
When should you choose Hailuo 2.3?
Choose Hailuo 2.3 when you need short video generation with both text and image inputs, especially when realistic motion, expressive characters, camera direction, or stylized rendering is more important than long duration. It is a sensible option for creative exploration, advertising concepts, image animation, social clips, and pre-production visuals where several short variations may be generated and edited together.
Consider another option when you need long-form or continuous video, native sound, documented function calling, fine-tuning, structured outputs, or a clearly supported current API model. Hailuo 2.3 Fast may be preferable for image-to-video batches where speed and cost are the main priorities. MiniMax H3 deserves comparison for new integrations because it is the newer primary video model in current developer documentation.
Overall, Hailuo 2.3 is best understood as a focused short-video model with a useful balance of image and text input support, cinematic motion targets, and predictable per-clip pricing. Its main trade-off is that the model is no longer the clearest forward-looking choice in MiniMax's catalog, and its short output limits make it a component for edited video workflows rather than a complete long-form production system.

