Grok Imagine

Grok Imagine Image

by xAI · Current and available through the xAI API as of September 24, 2026

xAI’s Grok Imagine Image is a dedicated API model for generating and editing images from text and image inputs. It supports 1K and 2K outputs, costs $0.02 per generated image plus $0.002 per input image, and is intended for visual applications rather than general text, coding, reasoning, speech, or video tasks.

Image generation Reasoning Coding
Grok Imagine Image is xAI’s specialized image-generation model for applications that need text-to-image creation or image editing through an API. Its canonical model ID is grok-imagine-image. The model accepts text prompts and image inputs, returns images at 1K or 2K resolution, and uses straightforward per-image pricing: $0.02 for each generated image and $0.002 for each input image.
Outputs

What Grok Imagine Image can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Batch API Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Grok Imagine
Model type Other
Context window 1K tokens
Status Current and available through the xAI API as of September 24, 2026
Knowledge cutoff notes

No model-specific knowledge cutoff is published in the current official image-model documentation. The model is an image-generation model, so a language-model training cutoff is not applicable as a directly documented specification.

Model notes

The canonical API model ID is grok-imagine-image. The model page lists grok-imagine-image-2026-03-02 as an alias; this alias is not a separate model entity. The API documentation lists version 1.0.0, a maximum prompt length of 1,024, 1K and 2K output resolutions, and a displayed rate limit of 6 requests per second. Batch API support is documented for us-east-1 and us-west-2, with batch pricing listed as N/A. The context_length value of 1024 represents the documented image-model prompt-length limit rather than a conventional language-model token context window. Editorial scores are not vendor benchmarks.

Cost

Model pricing

Input $0.002 per input image; text prompt pricing is not separately stated
Output $0.02 per generated image for 1K and 2K resolution
Model guide

Grok Imagine Image: xAI’s API Model for Image Generation and Editing

Grok Imagine Image is xAI’s dedicated API model for generating and editing images from text and image inputs. It produces 1K or 2K images, costs $0.02 per generated image plus $0.002 per input image, and is designed for visual applications rather than conversation, coding, reasoning, speech, video, or structured text generation.

What is Grok Imagine Image?

Grok Imagine Image is xAI’s current dedicated model for image generation and image editing. It is accessed through the xAI image-generation API using the canonical model ID grok-imagine-image. Unlike a general-purpose language model, it is designed specifically to produce visual output rather than conversational text.

The model supports two main workflows. A text prompt can describe a new image to generate, or a text prompt can be combined with an input image for editing and related image transformations. This makes it relevant to applications such as visual prototyping, creative tools, marketing graphics, and automated image-generation features.

Capabilities and supported modalities

Grok Imagine Image accepts both text and image inputs and produces image output. In practical terms, a request can contain a written description, an image to edit, or both. The provider documentation identifies image as the output modality.

  • Text input: Supported for describing the desired image or edit.
  • Image input: Supported for image editing and related transformations.
  • Image output: Supported at 1K and 2K resolution.
  • Audio and video: Not identified as supported input or output modalities for this model.
  • Text output: The model is not intended to return generated text as its primary result.

The documented maximum prompt length is 1,024. The available output sizes are 1K, defined in the documentation as 1024×1024, and 2K, defined as 2048×2048. The supplied specifications do not document additional aspect ratios, image formats, or output-count limits, so those details should not be assumed.

Where it fits in xAI’s lineup

Grok Imagine Image occupies the image-generation position in xAI’s current model catalog. It should not be confused with xAI’s separate video-generation offerings or with general-purpose Grok models intended for text interaction and reasoning. The model is listed with version 1.0.0, while grok-imagine-image-2026-03-02 is documented as an alias rather than a separate model entity.

As of September 24, 2026, the canonical model is listed as current and available through the xAI API. The documentation lists availability in the us-east-1 and us-west-2 regions.

Pricing and API availability

The provider’s documented pricing is based on images rather than on text tokens:

  • Generated image: $0.02 per image.
  • Input image: $0.002 per image.
  • Text prompt: No separate text-prompt charge is stated in the supplied documentation.
  • Resolution: The same documented output price applies to both 1K and 2K output.

For example, a request that edits one supplied image and generates one result would incur $0.002 for the input image and $0.02 for the generated image, for a documented image-related total of $0.022. Actual application costs can depend on the number of input and output images requested.

The Batch API is documented as supported in the listed regions, although batch pricing is shown as unavailable in the supplied research. The displayed rate limit is 6 requests per second. Developers should confirm current regional availability, limits, and billing details in xAI’s documentation before deploying a production integration.

Technical profile

SpecificationDocumented value
ProviderxAI
Canonical model IDgrok-imagine-image
Model version1.0.0
Input modalitiesText and image
Output modalityImage
Maximum prompt length1,024
Output resolutions1K (1024×1024) and 2K (2048×2048)
Generated-image price$0.02 per image
Input-image price$0.002 per image
Displayed rate limit6 requests per second
Batch APISupported in us-east-1 and us-west-2; batch pricing listed as unavailable

The 1,024 limit is described in the supplied model documentation as a maximum prompt length. It should not automatically be interpreted as a conventional language-model token context window, because this is an image model and the research does not define the unit as a standard text-token context length.

Strengths and trade-offs

The model’s clearest strength is specialization. It offers a direct API path for applications that need image generation or editing without selecting a general conversational model. The per-image price is also easy to understand, which can simplify budgeting for workloads where each request produces a predictable number of images.

Supporting image inputs makes it more useful than a text-only image generator for editing workflows. A product could, for example, pass an existing visual together with instructions describing a desired change. The documented 1K and 2K choices provide a basic quality-versus-output-size decision without introducing separate model IDs for each resolution.

Those advantages come with important limitations. Grok Imagine Image is not documented as a general-purpose reasoning or coding model. It does not provide text generation as its primary output, and the supplied specifications do not list tool use, function calling, streaming, fine-tuning, caching, structured text output, speech, transcription, or video generation. It therefore should not be selected for an application whose central task is conversation, code generation, data extraction, or multi-step text reasoning.

The model documentation shows a rate limit of 6 requests per second. That may be sufficient for moderate workloads, but high-volume systems must account for request scheduling, retries, and any limits that apply to the wider account or deployment. The supplied research also does not establish benchmark results, image-quality rankings, latency guarantees, or support for additional dimensions and image formats.

Speed, cost, and capability considerations

Editorially, the model is best understood as a relatively focused and cost-conscious image API rather than a broad AI system. The supplied database assigns it a speed score of 7 and a cost score of 8, but these are editorial evaluations, not xAI benchmarks or provider-published ratings. They should be treated as directional judgments rather than measured guarantees.

The $0.02 generated-image price can be attractive when an application needs many visual variations, provided that the requested number of images and any input-image charges are tracked carefully. Choosing 2K output does not add a separate documented per-image charge in the supplied pricing information, but it may still have practical implications for processing, storage, and delivery that are not quantified here.

For tasks where image quality, latency, or specialized editing behavior is the deciding factor, testing representative prompts and source images is more reliable than inferring performance from price alone. The available documentation establishes the billing and resolution options, but it does not publish comparative benchmark results against other image models.

Best use cases

Grok Imagine Image is a good fit when the central product requirement is API-based image creation or editing. Suitable examples include:

  • Text-to-image generation in creative applications.
  • Editing an existing image according to a written instruction.
  • Rapid visual prototyping during product or campaign development.
  • Generating concept visuals and marketing graphics.
  • Applications that need simple per-generated-image budgeting.
  • Workflows that need both prompt-based generation and image-conditioned editing through one image model.

Its image-only output also makes the expected result relatively clear: the application should be prepared to receive and process an image rather than a written explanation. If the surrounding workflow needs captions, structured metadata, code, or a conversation about the image, a separate text-capable model may be required; such an addition is outside the documented capabilities of Grok Imagine Image itself.

When to choose Grok Imagine Image

Choose Grok Imagine Image when you need a current xAI API model dedicated to generating or editing images, especially when straightforward per-image pricing and support for both text and image inputs matter. It is particularly appropriate for visual applications that do not need the model to reason through a long conversation or return structured text.

Another type of model may be more appropriate when the primary requirement is text generation, coding, tool calling, web interaction, speech, video, embeddings, or advanced multi-step reasoning. A different image option may also be preferable if your project requires capabilities not documented here, such as a specific aspect ratio, a particular image format, a published latency target, or independently verified quality benchmarks.

In short, Grok Imagine Image should be evaluated as a focused image-generation and editing component. Its main decision points are the text-and-image input support, 1K or 2K image output, $0.02 generated-image price, $0.002 input-image price, regional API availability, and the fact that it is not a general-purpose text model.

Current status and model identity

The canonical identifier is grok-imagine-image. The documented alias grok-imagine-image-2026-03-02 refers to the same model record and should not be counted as a separate model. The supplied research identifies the model as current and available through xAI’s API as of September 24, 2026. Because model documentation and pricing can change, production users should verify the current model page and pricing reference before implementation.


Answers to Frequently Asked Questions

Is Grok Imagine Image a general-purpose language model?
No. Grok Imagine Image is specialized for image generation and editing, with image as its primary output. It is not documented as a model for conversation, coding, tool use, structured text generation, speech, or video generation.
What is the model ID for Grok Imagine Image?
The canonical model ID is grok-imagine-image. The identifier grok-imagine-image-2026-03-02 is documented as an alias for the same model rather than a separate model.
How much does the Grok Imagine Image API cost?
A generated image costs $0.02, while an input image costs $0.002. For example, editing one supplied image and generating one result has a documented image-related cost of $0.022. The same listed generated-image price applies to both 1K and 2K output.
What inputs and outputs does Grok Imagine Image support?
The model accepts text and image inputs and produces image output. Text can describe a new image or an edit, while an input image can be supplied for editing and related transformations. Documented output resolutions are 1K (1024×1024) and 2K (2048×2048).
What is Grok Imagine Image used for?
Grok Imagine Image is xAI’s dedicated API model for generating new images from text prompts and editing existing images using text instructions. It is suitable for creative applications, visual prototyping, marketing graphics, and automated image-generation workflows.


Sources 5
Provider

About xAI