What is MiniMax T2V-01?
MiniMax T2V-01 is a proprietary text-to-video model developed by MiniMax for the Hailuo video-generation ecosystem. It was associated with MiniMax's first AI-native video-generation release, Video-01, announced on August 31, 2024. The T2V-01 designation describes the text-to-video version: the model accepts a written prompt and produces a video clip.
For example, a user could describe a cinematic shot, an animated scene, or a short visual concept in text and ask the model to render that description as video. Unlike an image-to-video system, T2V-01 was intended to begin from text rather than from an uploaded still image.
The model is not a general-purpose language model. It does not provide a conventional chat interface, text completion, coding assistance, reasoning workflow, or tool-calling system. Its primary output is generated video.
Verified output specifications
Historical first-party material and related documentation identify the following characteristics for T2V-01:
| Specification | Details |
|---|---|
| Primary input | Text prompt |
| Primary output | Video |
| Resolution | 720p |
| Frame rate | 25 frames per second |
| Maximum duration | Six seconds |
| Release timing | August 31, 2024 announcement |
These are historical specifications rather than a guarantee of what is available through a current MiniMax interface. The reviewed research does not verify a current endpoint, active quota, service-level commitment, or current price for this exact model.
How T2V-01 fits into MiniMax's video lineup
T2V-01 belongs to an early generation of MiniMax's Hailuo video models. It should be distinguished from T2V-01-Director, a later related model that added more explicit camera-control features and improved prompt adherence. The two names describe related but different model variants.
T2V-01 is also different from MiniMax image-to-video models. Text-to-video generation starts with a written description, while image-to-video generation uses an image as part of the source material. The supplied research does not verify image input for the standard T2V-01 model.
MiniMax's current official video-generation guide lists newer H3 and H3 Max models rather than T2V-01. This places T2V-01 in a legacy or superseded position within the provider's catalog. Its historical specifications remain useful for understanding the early Hailuo generation, but they should not be treated as evidence that the model is still generally available.
What T2V-01 was good at
T2V-01's main practical value was simple prompt-driven video creation. A user could describe a short visual sequence without preparing source footage or a reference image. That made the model suitable for early experimentation with AI-generated motion and for quickly testing whether a visual idea could be expressed as a short clip.
- Short concept clips: Six-second output is sufficient for testing a visual idea, scene, transition, or composition.
- Storyboarding: Short generated clips can help communicate the intended mood or direction of a scene before production work begins.
- Prompt experimentation: Text-only input lowers the preparation required to try different subjects, environments, and cinematic descriptions.
- Early Hailuo workflows: The model represents an early MiniMax approach to consumer-oriented text-to-video generation.
The 720p output and short duration also define its trade-off: the model was oriented toward compact creative experiments rather than long-form or high-resolution production.
Limitations and unsupported areas
The most important limitation is that T2V-01 is no longer presented as part of MiniMax's primary current video catalog. A user looking for a supported production model should therefore confirm availability before designing a workflow around it.
Its historical six-second maximum makes it unsuitable for projects that require continuous long scenes. A longer sequence would need to be assembled from multiple clips, and the supplied research does not verify any native continuation, editing, or shot-consistency feature for T2V-01.
The model's historical 720p output is also below the requirements of workflows that specifically need 1080p or higher resolution. The research does not verify native audio generation, audio synchronization, image or video reference inputs, structured output, function calling, streaming, fine-tuning, or a developer-facing context limit for this model.
Because T2V-01 is a video generator rather than a language model, conventional measures such as reasoning performance, coding ability, context-window size, and maximum text tokens are not applicable or were not published in the reviewed materials. It should be evaluated on visual generation needs, not on language-model benchmarks.
Pricing and access
A current model-specific price for MiniMax T2V-01 could not be verified. MiniMax has product-specific plans and credits across its services, but the supplied research does not provide a confirmed current price for this exact legacy model. Historical availability through Hailuo or related MiniMax services should not be assumed to mean that the same model, quota, or billing arrangement remains active today.
For anyone evaluating the model now, the practical first step is to check the current MiniMax video documentation or the relevant Hailuo interface. If T2V-01 is not listed, a newer supported model should be considered instead of relying on archived specifications or third-party listings.
When to choose this model
T2V-01 may be worth considering only when a user specifically needs to reproduce or study an early Hailuo text-to-video workflow and has verified that the model remains accessible in their account or service. Its narrow historical profile is clear: text prompt in, short 720p video out.
For a new project, choose a current MiniMax video model instead when active support, higher resolution, longer clips, newer generation quality, or additional controls matter. MiniMax's current documentation identifies H3 and H3 Max as the newer video-generation options. The supplied research does not provide a detailed feature-by-feature comparison, so exact differences in price, duration, resolution, and controls should be checked in the current documentation.
Another option may also be more appropriate if the workflow begins with a reference image, requires native audio, needs consistent characters across multiple shots, or depends on a stable production API. T2V-01's documented role does not establish support for those requirements.
Bottom line
MiniMax T2V-01 was an early Hailuo text-to-video model that generated six-second, 720p clips at 25 frames per second from text prompts. Its strengths were accessibility and fast visual experimentation, while its short duration, modest historical resolution, limited documented input mode, and uncertain current availability restrict its usefulness for modern production work. It is best treated as a legacy model and historical reference point within MiniMax's video-generation development, not as the default choice for a new workflow.

