Grok Imagine

Grok Imagine Image 2.0

by xAI · Generally available

xAI's Grok Imagine Image 2.0 is a dedicated image-generation and editing model. It accepts text and up to five reference images, supports configurable aspect ratios and 1K to 2K output tiers, and charges separately for input and generated images.

Image generation
Grok Imagine Image 2.0 is xAI's dedicated image model for text-to-image generation, natural-language editing, and multi-image compositing. Its main distinction is the combination of reference-image support, layout-sensitive generation, configurable output sizes, and a straightforward per-image API pricing model.
Outputs

What Grok Imagine Image 2.0 can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Batch API Multimodal output
Specifications

Technical details

Model family Grok Imagine
Model type Multimodal
Release date 2026-08-07
Status Generally available
Knowledge cutoff notes

No provider-published knowledge-cutoff date was found for this image-generation model. The model is documented as an image generation and editing endpoint rather than a token-based language model.

Model notes

Canonical API identifier is grok-imagine-image-2.0. The model accepts text and image inputs and supports up to five reference images for editing workflows. Documented output tiers include 1K, 1.5K, and 2K resolutions with Low and Medium quality options. Pricing is per image rather than token-based. Output prices are $0.04 for 1K Low, $0.05 for 1.5K Low, $0.06 for 2K Low, $0.06 for 1K Medium, $0.07 for 1.5K Medium, and $0.08 for 2K Medium. Batch API access is supported, but image generation uses standard pricing with no separate batch discount. Knowledge cutoff and token context limits are not applicable or not documented for this image-generation endpoint.

Cost

Model pricing

Input $0.01 per input image
Output $0.04-$0.08 per generated image, depending on resolution and quality
Model guide

Grok Imagine Image 2.0: Multi-Reference Image Generation and Editing

Grok Imagine Image 2.0 is xAI's image-generation and image-editing model for creating visuals from text, transforming existing images, and combining up to five reference images. It is available through the Imagine API and Grok products, with configurable 1K, 1.5K, and 2K output tiers priced per input and generated image.

What is Grok Imagine Image 2.0?

Grok Imagine Image 2.0 is an image-generation and image-editing model provided by xAI. It creates images from written prompts, but it can also use existing images as references for transformations, recomposition, and more controlled editing. The canonical API identifier is grok-imagine-image-2.0.

The model sits in xAI's Grok Imagine image-generation ecosystem rather than serving as a general-purpose conversational or reasoning model. It is available through the Imagine API and is also used by Grok's consumer-facing image-generation experience. That positioning makes it relevant both to people creating images interactively and to developers building image workflows into an application.

In practical terms, a user can describe a new scene, ask for changes to an existing image, or provide several reference images and request a combined result. The output is an image rather than text, code, audio, or video.

Core generation and editing capabilities

Image 2.0 supports text-to-image generation. A prompt can specify the subject, setting, visual style, composition, colors, typography, or intended format. The model is documented for use cases where composition and readable design elements matter, including marketing graphics, posters, banners, product imagery, illustrations, and concept art.

It also supports natural-language image editing. Instead of manually describing every pixel-level change, a user can provide an image and request an alteration such as changing the background, adjusting the composition, adding an object, or adapting the design for another format. The supplied research does not establish that every possible editing instruction will work equally well, so results should still be reviewed rather than treated as deterministic design operations.

A notable capability is support for up to five reference images in one request. Multiple references can be useful when a result needs to combine elements from different sources, such as a product from one image, a person or character from another, and a visual environment from a third. This makes the model suitable for multi-reference compositing and iterative visual development, although the final result may not preserve every reference detail exactly.

Output resolutions and aspect ratios

xAI documents configurable output resolution and aspect-ratio options for the model. The documented resolution tiers are 1K, 1.5K, and 2K, with Low and Medium quality options in the published pricing structure. Supported formats include square, portrait, landscape, and banner-oriented aspect ratios.

These choices are useful because the appropriate image shape depends on the destination. A square image may suit a product card or social post, a portrait layout may fit a mobile or poster design, and a wide banner can be used for headers or advertising placements. Selecting a larger resolution can improve the usefulness of an output for design work, but it also increases the per-image price.

The research does not provide pixel dimensions for every aspect-ratio and resolution combination, nor does it specify a maximum file size or a guaranteed output latency. Those details should be checked in xAI's current API documentation before building a workflow around them.

API pricing

Grok Imagine Image 2.0 uses image-based pricing rather than token pricing. Input images cost $0.01 per image. Generated output is priced separately according to resolution and quality:

Output optionPrice per generated image
1K Low$0.04
1.5K Low$0.05
2K Low$0.06
1K Medium$0.06
1.5K Medium$0.07
2K Medium$0.08

For a text-only generation request, there is no input-image charge, while an editing request using one reference image adds the $0.01 input-image charge. A request using five reference images would incur $0.05 in input-image charges before the generated-image fee. The exact total can therefore depend on both the number of supplied images and the selected output tier.

The model supports xAI's Batch API. However, the supplied pricing information indicates that image generation does not receive a separate batch discount. Batch processing may still be useful for organizing asynchronous workloads, but it should not be assumed to reduce the listed per-image price.

Supported modalities and technical scope

Image 2.0 accepts text and image inputs and produces image outputs. It is therefore multimodal in the specific sense that it can combine written instructions with visual references. Its primary output is a generated image.

  • Text input: Supported for prompts and editing instructions.
  • Image input: Supported, including up to five reference images in a request according to the supplied research.
  • Image output: Supported at documented 1K, 1.5K, and 2K tiers.
  • Audio and video input: Not documented for this model.
  • Audio and video output: Not supported as the model's output types.
  • Text generation: Not the purpose of this endpoint.
  • Tool or function calling: Not documented.

Context-window and maximum-token-output values do not apply in the usual language-model sense. xAI documents this model as an image-generation and editing endpoint rather than a token-based text model, and no token context limit or maximum output-token figure is supplied.

Reasoning, coding, and speed considerations

Grok Imagine Image 2.0 is not a reasoning model in the conversational sense. It does not provide a documented reasoning score, chain-of-thought capability, or general-purpose text analysis interface. It also does not provide coding capabilities or an execution environment. A developer can use code to call the API, but that is different from the model generating, reviewing, or running software.

The available research does not provide a verified latency target or speed rating. It would therefore be inaccurate to promise that Image 2.0 is faster than a particular competing image model. In general, the pricing structure creates a clear cost trade-off: 1K Low is the least expensive documented output option, while 2K Medium is the most expensive. Lower-cost settings may be appropriate for drafts and iteration; higher tiers are more suitable when the output needs additional resolution or the selected quality level.

Because image generation is probabilistic, producing several alternatives can cost more than generating a single image. Teams should account for exploratory generations, failed compositions, revisions, and input-image charges when estimating production costs.

Main strengths and limitations

Where the model is strongest

  • Combined generation and editing: The same model supports new images from prompts and natural-language changes to existing images.
  • Multi-reference workflows: Up to five reference images can support compositing and visual consistency tasks.
  • Design-oriented outputs: The documented use cases include typography, layouts, posters, banners, product imagery, and marketing assets.
  • Configurable formats: Resolution and aspect-ratio choices make it easier to target different destinations.
  • Predictable unit pricing: Costs are stated per input and generated image rather than being calculated from text-token usage.

Important limitations

  • No general assistant behavior: It is not intended for text conversations, general reasoning, coding, transcription, or speech generation.
  • No video generation: Video is outside this model's documented output scope.
  • No documented token limits: Language-model context and maximum-output-token specifications are not applicable or published for this endpoint.
  • No guaranteed visual fidelity: Reference images guide the result, but the supplied research does not promise exact preservation of identity, layout, text, or object details.
  • Usage and access can vary: Availability, rate limits, supported options, and access may depend on account configuration and API region.
  • Costs can accumulate during iteration: Multi-image requests and repeated generations each add to usage.

Best use cases

Grok Imagine Image 2.0 is a strong fit when the work is primarily visual and benefits from both prompt-based creation and image-guided editing. Suitable examples include:

  • Creating advertising concepts, social-media graphics, posters, and banners.
  • Generating early product-visualization concepts before a final photography or design pass.
  • Developing illustrations, concept art, and visual directions.
  • Combining several reference images into a new composition.
  • Adapting an existing visual to square, portrait, landscape, or banner formats.
  • Iterating on designs that contain layout elements or requested text.

For production work, it is sensible to separate exploration from final selection. Use lower-cost settings while testing concepts, then reserve higher-resolution or higher-quality outputs for candidates that have passed visual review.

When to choose Grok Imagine Image 2.0

Choose this model when you need a dedicated image endpoint with both text-to-image creation and multi-reference editing, especially if configurable 1K-to-2K output and per-image pricing fit your workflow. It is particularly relevant for developers building visual-generation features rather than a general chat assistant.

Another type of image model may be more appropriate when the priority is a specific capability not documented here, such as video creation, audio generation, exact production-grade design control, or a published latency guarantee. A language model is a better choice for coding, long-form text reasoning, or tool orchestration. If the main requirement is inexpensive experimentation, the 1K Low tier offers the lowest documented output price within Image 2.0; if the requirement is a larger final asset, the 1.5K or 2K options may be preferable despite their higher cost.

The most important trade-off is therefore specialization versus scope. Image 2.0 concentrates on image creation and editing instead of trying to handle conversation, software, speech, and video in one endpoint. That narrower scope can simplify an image pipeline, but it means additional models or tools are needed for non-visual tasks.

Bottom line

Grok Imagine Image 2.0 is xAI's image-focused model for generating and editing visuals with text and image inputs. Its defining practical features are support for up to five reference images, configurable aspect ratios and resolution tiers, and pricing that ranges from $0.04 to $0.08 per generated image, plus $0.01 for each input image. It is best evaluated as a visual-generation component rather than as a general-purpose AI model: useful for creative and design workflows, but not a substitute for text reasoning, coding, speech, or video systems.


Answers to Frequently Asked Questions

What is Grok Imagine Image 2.0 best used for?
The model is suited to advertising concepts, social-media graphics, posters, banners, product imagery, illustrations, concept art, multi-reference compositing, and adapting visuals to different formats. It is not intended for general conversation, coding, text reasoning, audio, or video generation.
What resolutions and aspect ratios does Grok Imagine Image 2.0 support?
The model supports documented 1K, 1.5K, and 2K resolution tiers, with Low and Medium quality options. Supported formats include square, portrait, landscape, and banner-oriented aspect ratios. Exact pixel dimensions for every combination should be confirmed in xAI’s current API documentation.
How much does the Grok Imagine Image 2.0 API cost?
Generated images cost between $0.04 and $0.08 each, depending on the selected resolution and quality. Documented prices range from $0.04 for 1K Low to $0.08 for 2K Medium. Input images cost an additional $0.01 per image, so a request with five reference images adds $0.05 before the output-generation fee.
What is Grok Imagine Image 2.0?
Grok Imagine Image 2.0 is xAI’s image-generation and image-editing model. It creates images from text prompts and can transform existing images using natural-language instructions and visual references. Its canonical API identifier is grok-imagine-image-2.0.
How many reference images can Grok Imagine Image 2.0 use?
Grok Imagine Image 2.0 supports up to five reference images in a single request. These images can be combined to guide elements such as products, characters, backgrounds, compositions, or visual styles, although the output may not preserve every reference detail exactly.


Sources 4
Provider

About xAI