Grok Imagine Video 1.5

Grok Imagine Video 1.5 Lite

by xAI · Current and available through the xAI Imagine API

xAI's Grok Imagine Video 1.5 Lite is a lower-cost video-generation model for text-to-video and image-to-video workflows. It supports 480p, 720p, and 1080p output, asynchronous API processing, Batch API access, and resolution-based per-second pricing.

Video generation Reasoning Coding
Grok Imagine Video 1.5 Lite is a specialized xAI video model designed for affordable, high-volume generation and rapid creative iteration. It accepts text and image inputs, produces video, supports 480p, 720p, and 1080p output, and uses per-second pricing through the xAI API.
Outputs

What Grok Imagine Video 1.5 Lite can produce

Video generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Batch API Multimodal output
Model profile

Performance characteristics

0/10 Reasoning
0/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Grok Imagine Video 1.5
Model type Lightweight
Context window tokens
Maximum output tokens
Status Current and available through the xAI Imagine API
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for this specialized video-generation model.

Model notes

Specialized video-generation model in xAI's Imagine API. The exact model page lists text and image inputs with video output, availability in us-east-1 and us-west-2, Batch API support, and a 10-requests-per-second limit. Video generation is asynchronous. Resolution-specific pricing is $0.02 per second at 480p, $0.03 per second at 720p, and $0.14 per second at 1080p. The model page does not publish a knowledge cutoff, context window, maximum output-token limit, release date, or fine-tuning option. A separate structured audio-output specification was not found for this exact Lite model.

Cost

Model pricing

Input $0.01 per image input; video input is charged per second where applicable
Output $0.02 per second at 480p; $0.03 per second at 720p; $0.14 per second at 1080p
Model guide

Grok Imagine Video 1.5 Lite: Affordable Text-to-Video and Image Animation

Grok Imagine Video 1.5 Lite is xAI's lower-cost video-generation model for creating short videos from text prompts or image references through the Imagine API.

What is Grok Imagine Video 1.5 Lite?

Grok Imagine Video 1.5 Lite is a specialized video-generation model from xAI. Its purpose is to turn a written description or an image reference into a short video clip. The model is positioned as the lower-cost Lite member of the Grok Imagine Video 1.5 family, making it more suitable for experimentation, high-volume drafts, and repeated creative revisions than workflows that always require the highest available output quality.

The canonical model identifier is grok-imagine-video-1.5-lite. It is available through xAI's Imagine API, rather than being a general-purpose conversational model. That distinction matters: the endpoint is built to generate video, not to answer questions, write code, perform extended reasoning, or return structured text.

xAI documents the model as available in the us-east-1 and us-west-2 regions. It also supports the Batch API, which can be useful when many independent video jobs need to be submitted for later processing.

How the model creates video

Grok Imagine Video 1.5 Lite supports two documented input patterns. A text-to-video request starts with a prompt describing the desired scene, subject, action, or visual direction. An image-to-video request uses an image as the starting reference and asks the model to animate or transform it.

Video generation is asynchronous. Instead of receiving the finished clip immediately, the API returns a request identifier. The application then polls the request until processing is complete and the video is available. This is an important implementation detail for user interfaces and production pipelines: applications should show progress or a pending state rather than treating the initial API response as the finished media.

The exact model documentation identifies text and image as inputs and video as the output. The supplied documentation does not establish a separate audio-generation capability, text-generation mode, image-output mode, or general-purpose tool-use interface for this model.

Resolutions and pricing

Grok Imagine Video 1.5 Lite supports 480p, 720p, and 1080p video output. xAI's documented output pricing is charged per generated second and varies substantially by resolution:

Output resolutionDocumented price
480p$0.02 per generated second
720p$0.03 per generated second
1080p$0.14 per generated second

The model listing also presents a $0.02-per-second headline output price, which corresponds to the documented 480p rate. For budgeting, the resolution-specific prices are more useful than the headline figure. A 10-second clip would therefore cost approximately $0.20 at 480p, $0.30 at 720p, or $1.40 at 1080p before any applicable input charges.

Image inputs are listed at $0.01 per image. The documentation also notes that video input can be charged per second where applicable, although the exact model capability listing identifies text and image inputs for this Lite model. Developers should confirm the current request format and any input-related charges before building a workflow around video inputs.

The resolution choice creates a clear quality-versus-cost trade-off. Lower resolution is better suited to rough drafts, prompt testing, and large batches. Higher resolution may be appropriate for a selected final clip, but the 1080p rate is much higher than the 480p and 720p rates, so generating every variation at 1080p can become expensive quickly.

Main strengths

Lower-cost iteration

The primary advantage of Grok Imagine Video 1.5 Lite is its cost positioning. At 480p, the per-second rate is low enough to support repeated prompt changes and multiple creative variations. This is useful when the first result is unlikely to be final and the main task is exploring motion, composition, scene direction, or visual style.

Text and image workflows

Supporting both text-to-video and image-to-video makes the model useful for more than purely prompt-driven generation. A creator can describe a scene from scratch, or begin with an existing still image and use the model to add movement. The latter can be useful for animating concept art, product images, illustrations, or other still visual assets, subject to the rights and permissions associated with those assets.

Multiple resolution choices

The availability of 480p, 720p, and 1080p allows a workflow to separate exploration from delivery. Teams can test ideas at a lower resolution and reserve 1080p generation for the small number of clips that have already been approved. This is more practical than treating every generation as a final-quality render.

Batch API access

Batch API support is another practical strength for high-volume work. It allows applications to organize many generation tasks without requiring every clip to be handled as an immediate interactive request. The model page lists a limit of 10 requests per second, so clients should still implement request management and avoid assuming unlimited submission capacity.

Limitations and undocumented specifications

Grok Imagine Video 1.5 Lite is a video specialist, not a general-purpose AI assistant. It should not be selected for chat, coding, document drafting, reasoning tasks, or structured text generation. The supplied model documentation does not publish a context window, maximum text-output limit, model-specific knowledge cutoff, or fine-tuning option. Those specifications are therefore unknown or not applicable to this video endpoint.

The model also does not provide a documented streaming video interface in the supplied research. Generation is asynchronous and may take several minutes depending on factors such as prompt complexity, duration, and resolution. Applications need to account for that delay, poll for completion, and handle temporary result URLs appropriately.

There is no documented reasoning score or coding capability for this model. It also has no listed function or tool-use support, and the endpoint should not be treated as a model that can browse the web, call external tools, or execute code. Its output is video rather than text, so it is not appropriate when the application needs a written explanation alongside or instead of the generated media.

Although the Lite model is designed for lower-cost generation, the supplied documentation does not publish a guaranteed generation-speed benchmark. In practice, request duration can vary, and the provider states that processing may take several minutes. The useful distinction is therefore operational rather than a fixed speed claim: Lite is intended for economical generation and iteration, while resolution, prompt complexity, clip duration, and queue conditions affect the actual wait.

Best use cases

  • Early-stage video experimentation: Generate several low-cost versions before choosing a direction.
  • Image animation: Turn a still image into a short moving clip using image-to-video input.
  • Social-media drafts: Produce short visual concepts where rapid variation matters more than maximum production quality.
  • Creative previsualization: Explore camera movement, scene action, and visual ideas before committing to more expensive rendering.
  • High-volume generation: Submit many independent jobs through the API or Batch API, while respecting the documented request limit.
  • Resolution-staged workflows: Test at 480p or 720p and use 1080p selectively for approved outputs.

When to choose Grok Imagine Video 1.5 Lite

Choose this model when the central requirement is affordable short-form video generation through an API. It is especially appropriate when a project needs many alternatives, supports asynchronous processing, and can use either a text prompt or a reference image. The 480p and 720p rates make it possible to keep exploration costs relatively controlled, while 1080p remains available for more demanding outputs.

A higher-cost video model may be more appropriate when maximum visual quality is more important than the cost of repeated generations. The supplied research specifically distinguishes this Lite model from the higher-cost Grok Imagine Video 1.5 model, but does not provide enough comparative specifications to claim exactly how their visual quality, speed, or motion handling differ. The safe distinction is that Lite is the lower-cost option, while the standard model is the higher-cost alternative in the same family.

A different type of model is preferable when the job requires text responses, coding, reasoning, web search, structured output, or interactive tool calls. Grok Imagine Video 1.5 Lite should be used as a media-generation component rather than as the central reasoning engine of an application.

Practical implementation checklist

  1. Use the exact model ID grok-imagine-video-1.5-lite.
  2. Choose 480p, 720p, or 1080p deliberately because pricing is resolution-dependent.
  3. Design the client for asynchronous processing and poll the returned request identifier.
  4. Allow for generation times that may extend to several minutes.
  5. Handle temporary result URLs according to the provider's current retention behavior.
  6. Use the Batch API for suitable high-volume workloads and stay within the documented 10-requests-per-second limit.
  7. Budget separately for image inputs and verify any applicable input charges before production deployment.

Overall, Grok Imagine Video 1.5 Lite is best understood as an economical, API-focused video generator. Its value comes from combining text and image inputs, several output resolutions, per-second pricing, and batch access. Its trade-offs are equally important: asynchronous generation, limited published specifications, no documented general-purpose reasoning or coding features, and a substantial price increase at 1080p.


Answers to Frequently Asked Questions

How does video generation work with the Grok Imagine Video 1.5 Lite API?
Video generation is asynchronous. The API returns a request identifier rather than the finished video immediately, so applications must poll the request until processing is complete. Generation may take several minutes, and the Batch API can be used for suitable high-volume workloads while respecting the documented limit of 10 requests per second.
Does Grok Imagine Video 1.5 Lite support text-to-video and image-to-video generation?
Yes. It can generate video from a written prompt or animate and transform a supplied image. The documented inputs are text and image, with video as the output; the supplied documentation does not establish separate audio-generation, image-output, or general-purpose text-generation capabilities.
What resolutions does Grok Imagine Video 1.5 Lite support?
The model supports 480p, 720p, and 1080p video output. Lower resolutions are generally more suitable for testing and high-volume iterations, while 1080p is better reserved for selected final clips because it costs significantly more.
What is Grok Imagine Video 1.5 Lite?
Grok Imagine Video 1.5 Lite is xAI's lower-cost video-generation model for creating short clips from text prompts or reference images. Its canonical model identifier is "grok-imagine-video-1.5-lite", and it is available through xAI's Imagine API.
How much does Grok Imagine Video 1.5 Lite cost?
Pricing depends on the output resolution: 480p costs $0.02 per generated second, 720p costs $0.03 per second, and 1080p costs $0.14 per second. Image inputs are listed at $0.01 per image. For example, a 10-second clip costs approximately $0.20 at 480p, $0.30 at 720p, or $1.40 at 1080p before applicable input charges.


Sources 4
Provider

About xAI