What is Grok Imagine Video 1.5?
Grok Imagine Video 1.5 is xAI's dedicated model for generating short videos. Instead of producing text or code, it creates video from a written prompt, a supplied still image, or a collection of reference images. It can also generate an audio track containing sound effects, ambience, and, where supported, dialogue.
The canonical API identifier is grok-imagine-video-1.5. xAI documentation also lists grok-imagine-video-1.5-preview and grok-imagine-video-1.5-2026-05-30 as preview or dated aliases rather than separate model entities. The model became generally available in the xAI API on June 16, 2026, according to the supplied release information.
Within xAI's current catalog, this is a specialized media-generation model rather than a general-purpose assistant. It is designed for visual creation workflows and is separate in purpose from models intended for reasoning, coding, web search, or structured text generation.
Supported generation workflows
Grok Imagine Video 1.5 supports three main ways to begin a video:
- Text-to-video: A written description is used to create a video. xAI describes this as a text-to-image step followed internally by image-to-video animation, although the API exposes it as one generation request.
- Image-to-video: A still image supplies the visual starting point, and the model animates it according to an accompanying instruction.
- Reference-to-video: Reference images guide the appearance or content of the generated clip. This workflow supports up to seven reference images and provides additional controls such as pinned first or last frames and keyframes.
Reference-to-video requests can also use up to three preset voices where supported. Custom uploaded voice references are restricted to trusted partners on request, so standard users should plan around the available preset voices rather than assuming arbitrary voice cloning is available.
Resolution, duration, and output handling
The model offers three documented output resolutions: 480p, 720p, and 1080p. Native 1080p is available for text-to-video and image-to-video generation. Reference-to-video is limited to 720p, making that workflow more flexible in terms of visual conditioning but less suitable when full-HD output is essential.
The documented maximum duration for reference-to-video is 15 seconds. Exact duration limits can vary by generation mode and current API constraints, so applications should not treat 15 seconds as a universal limit for every request type.
Completed videos are returned through temporary xAI-hosted URLs. Those URLs are useful for retrieving a result immediately, but they should not be treated as permanent storage. An application that needs long-term access should download the video and store it in its own durable storage after generation completes.
Generated audio and speech
Audio is a notable part of Grok Imagine Video 1.5's output. The model can generate sound effects, environmental ambience, and dialogue synchronized with the visual action. Audio generation is enabled by default, but applications can request a silent video when sound is unnecessary or will be added during post-production.
Preset voices are available for supported reference-to-video workflows, with up to three voices in a request according to the supplied documentation. The model therefore fits projects that need a quick audiovisual draft, such as a narrated concept, a social clip with ambient sound, or a product scene with synchronized effects. It should not be treated as a general audio-production system with unrestricted custom voice input.
Pricing by generated second
xAI prices Grok Imagine Video 1.5 according to the amount of video generated and the selected resolution:
| Output | Listed price |
|---|---|
| 480p video | $0.08 per generated second |
| 720p video | $0.14 per generated second |
| 1080p video | $0.25 per generated second |
| Image input | $0.01 per image |
| Preset audio input | Free |
For example, a 10-second 720p clip has a listed output charge of approximately $1.40 before any account-specific conditions. A 10-second 1080p clip would be approximately $2.50, while the same duration at 480p would be approximately $0.80. These calculations cover the listed generation charge and do not imply that other account, storage, or application costs are included.
Video generation is also available through xAI's Batch API. The supplied pricing information does not identify a separate batch discount for this model, so batch processing should be considered an execution option rather than assumed to reduce the per-second rate.
How the API works
Grok Imagine Video 1.5 uses an asynchronous generation process. An application starts a request, receives a request identifier, waits while the video is rendered, polls for completion, and then retrieves the temporary video URL. This differs from a typical text-generation call that returns a response immediately.
The xAI SDK provides a higher-level video-generation interface that can handle polling. Developers using REST directly must implement the request, status checks, completion handling, error handling, and download or storage steps themselves. The model is documented as available in the us-east-1 and us-west-2 regions.
Because the model's output is video with optional audio, it does not provide text-model features such as reasoning scores, coding generation, function calling, structured JSON output, or web search. The supplied research also does not specify a context window or maximum token output, which are not the primary limits for this media-generation model.
Main strengths
- Several creative entry points: Text, still images, and multiple reference images can all drive generation.
- Native 1080p for key workflows: Text-to-video and image-to-video can produce 1080p output without relying on an external upscaling step.
- Audio in the same generation pass: Sound effects, ambience, and supported dialogue can be synchronized with the video.
- Useful reference controls: Up to seven reference images, pinned frames, keyframes, and preset voices give reference-based workflows more control than a basic text-only prompt.
- Predictable output pricing: The per-second rates make it possible to estimate the direct generation cost before rendering a clip.
Limitations and trade-offs
The main cost trade-off is resolution. Moving from 480p to 1080p increases the listed per-second rate substantially, so high-resolution experimentation can become expensive when many variations are required. A practical workflow may use lower resolution for early prompt iteration and reserve 1080p for selected final candidates.
Reference-to-video has a separate trade-off: it offers more visual guidance but is capped at 720p. Users who need detailed 1080p animation from a supplied image should consider image-to-video instead when the desired workflow allows it.
The model is also not a replacement for a deterministic editing system. The supplied specifications do not promise frame-perfect editing, long-form production, or precise timeline manipulation. Temporary result URLs add another operational requirement, since production applications need to copy outputs to durable storage.
Finally, Grok Imagine Video 1.5 is specialized for media generation. It is not the appropriate choice for general-purpose reasoning, coding, data extraction, structured-output tasks, or tool-driven assistants. Those jobs belong with a different model type in xAI's broader ecosystem.
When to choose Grok Imagine Video 1.5
Choose this model when the central requirement is short-form video creation and you want several ways to control the visual result. It is particularly suitable for:
- Rapid creative prototypes and storyboards.
- Short marketing or social-media clips.
- Animating product images or cinematic stills.
- Visual concept development for advertisements, games, or film projects.
- Applications that need generated sound effects, ambience, or supported dialogue alongside the video.
- Automated services that can work with an asynchronous generation-and-download pipeline.
It may be a poor fit when the project requires long-form video, precise frame-level edits, unrestricted custom voice references, or large volumes of inexpensive 1080p generations. It is also the wrong choice when video is only a secondary feature and the primary task is reasoning, coding, browsing, or structured text production.
Bottom line
Grok Imagine Video 1.5 is best understood as a focused video-generation service rather than a general AI model. Its strongest combination is support for text-to-video, image-to-video, and reference-guided workflows, plus native 1080p in the first two modes and optional generated audio. The most important practical constraints are resolution-dependent pricing, the 720p limit for reference-to-video, asynchronous processing, temporary output URLs, and the lack of general-purpose language or tool-use features.
For creators and developers who need short audiovisual clips and can manage per-second costs and asynchronous delivery, it offers a broad set of generation controls in one API. For precise editing, long productions, or non-media tasks, a dedicated editing pipeline or another model category will be more appropriate.

