What is Grok Imagine Image?
Grok Imagine Image is xAI’s current dedicated model for image generation and image editing. It is accessed through the xAI image-generation API using the canonical model ID grok-imagine-image. Unlike a general-purpose language model, it is designed specifically to produce visual output rather than conversational text.
The model supports two main workflows. A text prompt can describe a new image to generate, or a text prompt can be combined with an input image for editing and related image transformations. This makes it relevant to applications such as visual prototyping, creative tools, marketing graphics, and automated image-generation features.
Capabilities and supported modalities
Grok Imagine Image accepts both text and image inputs and produces image output. In practical terms, a request can contain a written description, an image to edit, or both. The provider documentation identifies image as the output modality.
- Text input: Supported for describing the desired image or edit.
- Image input: Supported for image editing and related transformations.
- Image output: Supported at 1K and 2K resolution.
- Audio and video: Not identified as supported input or output modalities for this model.
- Text output: The model is not intended to return generated text as its primary result.
The documented maximum prompt length is 1,024. The available output sizes are 1K, defined in the documentation as 1024×1024, and 2K, defined as 2048×2048. The supplied specifications do not document additional aspect ratios, image formats, or output-count limits, so those details should not be assumed.
Where it fits in xAI’s lineup
Grok Imagine Image occupies the image-generation position in xAI’s current model catalog. It should not be confused with xAI’s separate video-generation offerings or with general-purpose Grok models intended for text interaction and reasoning. The model is listed with version 1.0.0, while grok-imagine-image-2026-03-02 is documented as an alias rather than a separate model entity.
As of September 24, 2026, the canonical model is listed as current and available through the xAI API. The documentation lists availability in the us-east-1 and us-west-2 regions.
Pricing and API availability
The provider’s documented pricing is based on images rather than on text tokens:
- Generated image: $0.02 per image.
- Input image: $0.002 per image.
- Text prompt: No separate text-prompt charge is stated in the supplied documentation.
- Resolution: The same documented output price applies to both 1K and 2K output.
For example, a request that edits one supplied image and generates one result would incur $0.002 for the input image and $0.02 for the generated image, for a documented image-related total of $0.022. Actual application costs can depend on the number of input and output images requested.
The Batch API is documented as supported in the listed regions, although batch pricing is shown as unavailable in the supplied research. The displayed rate limit is 6 requests per second. Developers should confirm current regional availability, limits, and billing details in xAI’s documentation before deploying a production integration.
Technical profile
| Specification | Documented value |
|---|---|
| Provider | xAI |
| Canonical model ID | grok-imagine-image |
| Model version | 1.0.0 |
| Input modalities | Text and image |
| Output modality | Image |
| Maximum prompt length | 1,024 |
| Output resolutions | 1K (1024×1024) and 2K (2048×2048) |
| Generated-image price | $0.02 per image |
| Input-image price | $0.002 per image |
| Displayed rate limit | 6 requests per second |
| Batch API | Supported in us-east-1 and us-west-2; batch pricing listed as unavailable |
The 1,024 limit is described in the supplied model documentation as a maximum prompt length. It should not automatically be interpreted as a conventional language-model token context window, because this is an image model and the research does not define the unit as a standard text-token context length.
Strengths and trade-offs
The model’s clearest strength is specialization. It offers a direct API path for applications that need image generation or editing without selecting a general conversational model. The per-image price is also easy to understand, which can simplify budgeting for workloads where each request produces a predictable number of images.
Supporting image inputs makes it more useful than a text-only image generator for editing workflows. A product could, for example, pass an existing visual together with instructions describing a desired change. The documented 1K and 2K choices provide a basic quality-versus-output-size decision without introducing separate model IDs for each resolution.
Those advantages come with important limitations. Grok Imagine Image is not documented as a general-purpose reasoning or coding model. It does not provide text generation as its primary output, and the supplied specifications do not list tool use, function calling, streaming, fine-tuning, caching, structured text output, speech, transcription, or video generation. It therefore should not be selected for an application whose central task is conversation, code generation, data extraction, or multi-step text reasoning.
The model documentation shows a rate limit of 6 requests per second. That may be sufficient for moderate workloads, but high-volume systems must account for request scheduling, retries, and any limits that apply to the wider account or deployment. The supplied research also does not establish benchmark results, image-quality rankings, latency guarantees, or support for additional dimensions and image formats.
Speed, cost, and capability considerations
Editorially, the model is best understood as a relatively focused and cost-conscious image API rather than a broad AI system. The supplied database assigns it a speed score of 7 and a cost score of 8, but these are editorial evaluations, not xAI benchmarks or provider-published ratings. They should be treated as directional judgments rather than measured guarantees.
The $0.02 generated-image price can be attractive when an application needs many visual variations, provided that the requested number of images and any input-image charges are tracked carefully. Choosing 2K output does not add a separate documented per-image charge in the supplied pricing information, but it may still have practical implications for processing, storage, and delivery that are not quantified here.
For tasks where image quality, latency, or specialized editing behavior is the deciding factor, testing representative prompts and source images is more reliable than inferring performance from price alone. The available documentation establishes the billing and resolution options, but it does not publish comparative benchmark results against other image models.
Best use cases
Grok Imagine Image is a good fit when the central product requirement is API-based image creation or editing. Suitable examples include:
- Text-to-image generation in creative applications.
- Editing an existing image according to a written instruction.
- Rapid visual prototyping during product or campaign development.
- Generating concept visuals and marketing graphics.
- Applications that need simple per-generated-image budgeting.
- Workflows that need both prompt-based generation and image-conditioned editing through one image model.
Its image-only output also makes the expected result relatively clear: the application should be prepared to receive and process an image rather than a written explanation. If the surrounding workflow needs captions, structured metadata, code, or a conversation about the image, a separate text-capable model may be required; such an addition is outside the documented capabilities of Grok Imagine Image itself.
When to choose Grok Imagine Image
Choose Grok Imagine Image when you need a current xAI API model dedicated to generating or editing images, especially when straightforward per-image pricing and support for both text and image inputs matter. It is particularly appropriate for visual applications that do not need the model to reason through a long conversation or return structured text.
Another type of model may be more appropriate when the primary requirement is text generation, coding, tool calling, web interaction, speech, video, embeddings, or advanced multi-step reasoning. A different image option may also be preferable if your project requires capabilities not documented here, such as a specific aspect ratio, a particular image format, a published latency target, or independently verified quality benchmarks.
In short, Grok Imagine Image should be evaluated as a focused image-generation and editing component. Its main decision points are the text-and-image input support, 1K or 2K image output, $0.02 generated-image price, $0.002 input-image price, regional API availability, and the fact that it is not a general-purpose text model.
Current status and model identity
The canonical identifier is grok-imagine-image. The documented alias grok-imagine-image-2026-03-02 refers to the same model record and should not be counted as a separate model. The supplied research identifies the model as current and available through xAI’s API as of September 24, 2026. Because model documentation and pricing can change, production users should verify the current model page and pricing reference before implementation.
Answers to Frequently Asked Questions
grok-imagine-image. The identifier grok-imagine-image-2026-03-02 is documented as an alias for the same model rather than a separate model.
