What is MiniMax T2V-01-Director?
MiniMax T2V-01-Director is a text-to-video model from MiniMax's Hailuo video ecosystem. It was announced on March 3, 2025, alongside the image-to-video model I2V-01-Director. The two models serve different workflows: T2V-01-Director starts with a written prompt, while I2V-01-Director is intended for animating an image.
The defining feature of T2V-01-Director is directed camera movement. Instead of describing only a subject and setting, users can specify how the virtual camera should behave during the shot. This makes the model more relevant to storyboards, advertising concepts, visual development, and short cinematic experiments than a basic text-to-video system that offers little control over shot composition.
MiniMax announced the model as globally available through Hailuo Video and the MiniMax Open Platform. However, its current catalog position requires caution: the current official pricing material reviewed for this record prominently lists newer MiniMax video models and does not provide a verified current price for T2V-01-Director. It is therefore best treated as a legacy or potentially limited-availability model until access is confirmed in the relevant Hailuo or developer interface.
How the camera-control feature works
A conventional text-to-video prompt might describe a person walking through a city at night. A T2V-01-Director prompt can add instructions such as a slow push toward the subject, a lateral tracking shot, or a tilt from the street to a building. These instructions give the generation process a shot-level direction in addition to the scene description.
MiniMax's launch material describes 15 available camera movements that can be combined. The documented movement categories include operations such as pans, tilts, pushes, pulls, and tracking shots. The research does not establish that every possible combination will be equally reliable, so camera control should be understood as a directing aid rather than a guarantee of exact physical cinematography.
A practical prompt can combine four elements:
- Subject: what appears in the shot, such as a cyclist, product, character, or landscape.
- Environment: the location, time of day, lighting, and surrounding details.
- Action: what the subject does during the clip.
- Camera direction: the movement, framing change, or transition the virtual camera should perform.
For example, a creator might describe a red car driving through a rain-soaked street at night, then add a slow tracking movement that follows the vehicle from the side. The model's value is not simply that it generates motion; it attempts to make the intended shot language part of the generation request.
Verified technical profile
The following specifications are supported by the supplied research or by the model record. Some values are not applicable to this type of video model and should not be inferred from language-model conventions.
| Attribute | Documented information |
|---|---|
| Provider | MiniMax |
| Model family | Video-01 |
| Model type | Text-to-video |
| Primary input | Text prompts |
| Primary output | Generated video |
| Documented resolution | 720P |
| Documented clip duration | 6 seconds per clip |
| Camera control | 15 documented camera movements and combinations, according to MiniMax launch material |
| Release date | March 3, 2025 |
| Image input | Not documented for this model; image-to-video is handled by I2V-01-Director |
| Text output | Not supported as a model output |
| Audio output | Not verified |
The six-second duration and 720P resolution are reported in the supplied model documentation and secondary listings, while the official launch announcement establishes the model's Director-series positioning and camera-control feature. The current record does not provide a verified context window, maximum token output, image-input limit, audio specification, or current API quota.
Input, output, and capabilities
T2V-01-Director accepts text and produces video. Its multimodal classification reflects video output rather than a general multimodal assistant interface. The supplied research does not verify image input, video input, audio input, text generation, image generation, speech, music, embeddings, or structured JSON output for this specific model.
The model is also not documented as a reasoning or coding model. It may interpret scene descriptions and camera instructions, but that should not be confused with a dedicated reasoning capability. It is not an appropriate choice for code generation, analysis, chat, document processing, or other language-centered tasks.
No verified support is established for web search, tool or function calling, streaming, fine-tuning, caching, batch processing, or structured outputs. These omissions matter for production workflows. A user who needs a predictable machine-readable response, a multi-step tool-using agent, or programmatic control beyond video generation should confirm whether a newer MiniMax video service supplies those features rather than assuming they are inherited from the wider MiniMax platform.
Where the model is strongest
T2V-01-Director's main strength is directorial control over short shots. Many text-to-video workflows focus primarily on subject appearance and broad motion. This model gives the prompt a more explicit cinematographic role, which can help users plan a shot around camera movement rather than repeatedly regenerating until an acceptable movement appears.
- Shot planning: Camera instructions can make generated clips more useful as visual storyboards.
- Advertising concepts: Product or lifestyle concepts can be framed with a specified push, pan, tilt, or tracking movement.
- Short-form cinematic content: Six-second clips fit quick visual experiments and social-media-oriented concept development.
- Animation and previsualization: Creators can test how a scene might look before committing to a full production.
- Controlled experimentation: Combining camera directions can provide a more deliberate alternative to unconstrained prompt-only generation.
These are practical use cases based on the model's documented purpose, not benchmark claims. The supplied research contains no authoritative quality benchmark, consistency score, or comparison showing that it outperforms newer video models on general visual fidelity.
Limitations and availability concerns
The model is designed for short clips, not long-form video production. A six-second maximum documented duration means that longer sequences would require multiple generations and external editing. That can introduce continuity problems in subjects, environments, lighting, and camera position, although the supplied research does not quantify how often those problems occur.
The fixed 720P output is suitable for previews, concepts, and some short-form uses, but it may be insufficient when a project requires higher-resolution delivery. The model also does not replace an image-to-video workflow. If the starting point is a particular illustration, photograph, or designed frame, the separate I2V-01-Director model is the more relevant member of the Director series.
Current access is another limitation. MiniMax's reviewed official pricing page lists newer video models but does not provide a verified current price for T2V-01-Director. No current input or output price is available in the supplied research. Users should check the Hailuo interface or MiniMax developer console directly before planning a paid workflow. Availability, quotas, and model access may have changed since the 2025 launch.
The model record classifies T2V-01-Director as legacy or limited availability, but no deprecation or shutdown date has been verified. That distinction is important: the model may still be accessible in some environment without being a prominently supported option in the current catalog.
Pricing and API status
No verified current price is available for T2V-01-Director. The supplied research specifically avoids assigning a token price because token-oriented limits and pricing are not documented for this video-generation model. Third-party listings may describe historical or reseller access, but they should not be treated as MiniMax's current first-party pricing.
The model was announced for the MiniMax Open Platform as well as Hailuo Video, but the current record does not confirm that first-party API access remains available, nor does it establish request syntax, authentication requirements, quotas, rate limits, or production service-level commitments. Developers should verify all of these details in the current MiniMax console before building an integration.
When to choose T2V-01-Director
Choose T2V-01-Director when the specific requirement is a short text-generated video with explicit camera-direction control and when the model is available in your chosen MiniMax environment. It is a sensible candidate for a six-second concept shot, a storyboard panel with motion, a product-advertising draft, or a visual test where the camera move matters as much as the subject.
Its strongest case is not maximum resolution, long duration, or broad AI assistance. It is the combination of text-to-video generation and a defined set of cinematic movements. That specialization can be more useful than a newer general video model when a creator needs to communicate a particular shot idea quickly.
Another option may be more appropriate when the project needs image-to-video animation, long scenes, higher output resolution, synchronized audio, reliable structured API responses, or a currently documented price and production interface. For an image-led workflow, I2V-01-Director is the directly related alternative identified in the launch material. For a current production project, newer models in MiniMax's catalog may also be worth evaluating, but the supplied research does not establish a feature-by-feature comparison or guarantee that any particular successor preserves the same camera-control behavior.
Bottom line
MiniMax T2V-01-Director is a focused text-to-video model whose practical distinction is cinematic camera control. It generates documented six-second, 720P clips and supports a set of camera movements that can be combined with descriptions of subjects, environments, action, and style. That makes it useful for short visual concepts and previsualization.
At the same time, it should not be evaluated like a general-purpose AI model. It has no verified language, coding, search, tool-calling, structured-output, or audio-generation role, and its current price and catalog availability are not confirmed. Before relying on it for production, verify access and limits directly with MiniMax. For the right short-shot workflow, however, its camera-direction focus remains the reason to consider it.

