GPT Image

GPT-Image-2

by OpenAI · Active; GPT Image 2.5 models are available for newer workflows, but GPT-Image-2 remains accessible as a documented API model.

OpenAI’s GPT-Image-2 generates and edits images from text prompts and reference images. It supports high-fidelity image inputs, custom resolutions, improved text rendering, PNG, JPEG, and WebP outputs, transparent backgrounds in preview, and discounted Batch API processing. Token-based pricing and the lack of text, audio, video, function-calling, structured-output, and streaming support make it best suited to dedicated image workflows.

Image generation
GPT-Image-2 is OpenAI’s dedicated model for generating and editing images through the Images API. It can turn written descriptions into images, revise existing images, and use reference images while preserving important visual details. The model supports custom resolutions, improved text rendering, PNG, JPEG, and WebP output, and high-fidelity image input processing. It is a focused visual model rather than a conversational or reasoning system, so its value depends on whether an application needs reliable image creation and editing instead of text generation, coding, audio, or video.
Outputs

What GPT-Image-2 can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Prompt caching Batch API Multimodal output
Model profile

Performance characteristics

8/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family GPT Image
Model type Multimodal
Release date 2026-04-21
Status Active; GPT Image 2.5 models are available for newer workflows, but GPT-Image-2 remains accessible as a documented API model.
Knowledge cutoff notes

OpenAI’s public GPT-Image-2 documentation does not specify a knowledge cutoff for this image-generation model.

Model notes

Canonical model ID is gpt-image-2. A dated snapshot, gpt-image-2-2026-04-21, is also available. GPT-Image-2 accepts text and image inputs and produces image output. Image inputs are always processed at high fidelity, so input_fidelity should be omitted. Supported quality values are low, medium, high, and auto. Custom resolutions must use dimensions divisible by 16, remain within the documented pixel and aspect-ratio limits, and not exceed a 3,840-pixel edge. PNG, JPEG, and WebP outputs are supported. Transparent backgrounds are available in preview for PNG and WebP. Batch API processing is supported with a 50% discount. Pricing is token-based and varies with prompt length, reference-image inputs, quality, and resolution.

Cost

Model pricing

Input $8.00 per 1M image input tokens; $2.00 per 1M cached image input tokens; $5.00 per 1M text input tokens; $1.25 per 1M cached text input tokens
Output $30.00 per 1M image output tokens; $10.00 per 1M text output tokens where applicable
Model guide

GPT-Image-2: Features, Pricing, Capabilities and API Support

GPT-Image-2 is OpenAI’s image generation and editing model for creating visuals from text prompts and modifying reference images. It supports high-fidelity image inputs, flexible resolutions, improved text rendering, multiple output formats, transparent backgrounds in preview, and Batch API processing. Its token-based pricing and lack of text, audio, video, function-calling, structured-output, and streaming support make it best suited to production image workflows rather than general-purpose AI applications.

What is GPT-Image-2?

GPT-Image-2 is OpenAI’s image generation and editing model. It accepts text prompts and image inputs, then produces generated or edited images. The canonical API model identifier is gpt-image-2, and OpenAI also provides the dated snapshot gpt-image-2-2026-04-21 for applications that need a fixed model version.

The model is designed for workflows where the visual result matters more than producing a written answer. For example, a developer can use it to create a marketing illustration from a description, revise a product image using a reference picture, or generate an editorial graphic containing readable text. GPT-Image-2 is available through OpenAI’s Images API and also supports Batch API processing.

Where GPT-Image-2 fits in OpenAI’s lineup

GPT-Image-2 belongs to OpenAI’s GPT Image family and is focused specifically on visual generation and editing. It should not be treated as a smaller version of a general-purpose GPT model. Its primary output is an image, and the supplied specifications do not describe it as a text reasoning, coding, speech, or video model.

OpenAI’s current catalog also includes newer GPT Image 2.5 models for some workflows. The available research does not provide a complete feature or price comparison for those models, so GPT-Image-2 should be evaluated on its documented capabilities rather than assumed to be superior or inferior in every situation. Readers comparing the product family can also review GPT-Image-2.5 Sunburst and GPT-Image-2.5 Flare where those alternatives are relevant.

Supported inputs and outputs

GPT-Image-2 supports text input and image input. Image input is useful for editing, restyling, preserving reference details, or building a new composition around an existing visual. The model’s output is an image rather than a text response.

  • Text input: Supported for describing the desired image or edit.
  • Image input: Supported for reference-based generation and editing.
  • Image output: Supported as the model’s primary result.
  • Audio and video input: Not supported.
  • Audio and video output: Not supported.
  • Text, embedding, and structured machine-readable output: Not provided as native model outputs.

Image inputs are always processed at high fidelity. As a result, applications using GPT-Image-2 should omit the input_fidelity parameter rather than trying to select a lower or higher fidelity mode.

Resolution, quality and output formats

The model supports automatic sizing as well as custom resolutions. Custom dimensions must be multiples of 16. The longest edge cannot exceed 3,840 pixels, the aspect ratio cannot exceed 3:1, and the image must remain within OpenAI’s documented total-pixel limits. Resolutions above 2K are considered experimental, meaning that results can be more variable at larger sizes.

Generation quality can be set to low, medium, high, or auto. This gives developers a practical quality-versus-cost and quality-versus-speed control, although the supplied documentation does not provide a fixed generation time for each setting. Higher requested quality and larger images can also affect token consumption.

GPT-Image-2 supports PNG, JPEG, and WebP output. Transparent backgrounds are available in preview when using PNG or WebP. These format and transparency options are useful for product cutouts, interface assets, logos, stickers, composited marketing graphics, and other images that may need to be placed over a separate background.

GPT-Image-2 pricing and Batch API support

GPT-Image-2 uses token-based pricing rather than a single fixed price per generated image. The amount consumed depends on factors such as the prompt, reference-image inputs, requested resolution, and quality setting.

Usage typePublished price
Image input$8 per 1 million tokens
Cached image input$2 per 1 million tokens
Image output$30 per 1 million tokens
Text input$5 per 1 million tokens
Cached text input$1.25 per 1 million tokens
Text output, where applicable$10 per 1 million tokens

These rates do not translate directly into a universal per-image price because image dimensions, quality, prompts, and reference images change token usage. Developers should estimate costs using the actual workloads they expect to send rather than assuming every image has the same price.

Batch API requests receive a 50% discount according to OpenAI’s release documentation. Batch processing is therefore more appropriate for queued or non-interactive workloads, such as preparing a large set of catalog images or editorial assets, than for an application that must return an image immediately.

Main strengths and limitations

Strengths

  • Reference-image handling: High-fidelity image input processing helps workflows that need to preserve details from supplied images.
  • Text inside images: OpenAI describes improved text rendering, which is relevant to posters, labels, diagrams, advertisements, and other text-heavy visual assets.
  • Flexible layouts: Custom resolutions and a broad aspect-ratio limit support outputs beyond standard square, portrait, and landscape presets.
  • Production formats: PNG, JPEG, and WebP make the output practical for different delivery and editing requirements.
  • Transparent backgrounds: Preview support for PNG and WebP can simplify compositing and asset preparation.
  • Batch workflows: Batch API support and its discount can reduce the cost of large, asynchronous jobs.

Limitations

  • GPT-Image-2 is not a general-purpose conversational model and does not provide native text reasoning or coding output.
  • It does not support audio or video input or output.
  • Function calling, structured outputs, and streaming are not supported according to the supplied model specifications.
  • There is no documented context-window or maximum-output-token value for this image model.
  • Pricing varies with token usage instead of following a simple fixed price per image.
  • Resolutions above 2K are experimental, and custom image dimensions must follow strict multiple-of-16, pixel, edge, and aspect-ratio constraints.

Reasoning, coding and tool support

GPT-Image-2 should not be selected for tasks that require a reasoning model to analyze a problem and return a detailed text answer. The supplied specifications do not assign it a reasoning score or coding capability, and its text output is not a native model output. It is also not documented as supporting function calling, web search, or other tool use.

An application can still combine GPT-Image-2 with a separate language model or application layer. For example, another component could interpret a user’s request and produce an image prompt, while GPT-Image-2 handles the visual generation. That architecture does not make GPT-Image-2 itself a text reasoning or tool-using model; the responsibilities belong to different components.

Speed and cost trade-offs

The available editorial assessment rates GPT-Image-2’s speed at 8 out of 10 and cost at 6 out of 10. These are comparative editorial scores, not OpenAI-published benchmarks or guarantees. The documented controls provide a more concrete way to understand the trade-off: developers can choose low, medium, high, or automatic quality, while output size and image complexity influence token consumption.

For interactive applications, lower quality and moderate resolutions may be more appropriate when users need quick previews. High-quality, large-format images are better suited to final assets but may consume more tokens and produce less predictable latency. Batch processing is a separate option for workloads where immediate responses are unnecessary and the 50% discount is more valuable than real-time delivery.

Best use cases for GPT-Image-2

GPT-Image-2 is a strong fit for developers building image-focused products and workflows such as:

  • Text-to-image creation tools for marketing, editorial, or creative teams.
  • Reference-based product visualization and image revision.
  • Advertising assets that require readable text within the image.
  • Flexible-format social, web, packaging, or campaign graphics.
  • Transparent product cutouts, stickers, icons, and composited design elements.
  • Batch generation of catalog, campaign, or publishing assets.
  • Applications where preserving important details from a supplied image is more important than generating an unrelated visual.

When to choose GPT-Image-2

Choose GPT-Image-2 when the central requirement is high-quality image generation or editing, especially when the workflow needs reference images, custom dimensions, readable embedded text, multiple output formats, or transparent backgrounds. It is also a reasonable choice when an application can manage token-based billing and does not need a fixed per-image price.

Consider another type of model when the main task is conversation, document analysis, coding, speech, video creation, structured JSON generation, or tool orchestration. A general-purpose text model is more appropriate for reasoning and code, while a video or audio model is needed for those media types. If the application requires streaming responses, function calling, or deterministic structured output, GPT-Image-2 is not the right standalone component.

GPT-Image-2 is therefore best understood as a specialized visual engine. Its advantage is control over image creation and editing rather than breadth across every modality. The most suitable alternative depends on whether the priority is lower cost, simpler fixed-image pricing, faster previews, richer text reasoning, or support for a different media format.


Answers to Frequently Asked Questions

Does GPT-Image-2 support Batch API processing?
Yes. GPT-Image-2 supports Batch API processing, which is suitable for queued, non-interactive workloads such as generating catalog or editorial assets. Batch requests receive a 50% discount according to OpenAI’s release documentation, but they are not intended for applications that require immediate image responses.
What image formats, resolutions, and quality settings does GPT-Image-2 support?
GPT-Image-2 supports PNG, JPEG, and WebP outputs, with transparent backgrounds available in preview for PNG and WebP. It offers low, medium, high, and auto quality settings, supports custom dimensions that are multiples of 16, and allows a longest edge of up to 3,840 pixels subject to documented pixel and aspect-ratio limits. Resolutions above 2K are experimental.
How much does GPT-Image-2 cost?
GPT-Image-2 uses token-based pricing. Published rates include $8 per 1 million image-input tokens, $2 for cached image input, $30 for image-output tokens, $5 for text input, $1.25 for cached text input, and $10 for applicable text output. The actual cost per image varies with resolution, quality, prompts, and reference-image usage.
What is GPT-Image-2 used for?
GPT-Image-2 is OpenAI’s image generation and editing model. It can create images from text prompts, edit or restyle images using reference inputs, generate visuals with readable embedded text, and produce assets for marketing, editorial, product, social media, and design workflows.
What inputs and outputs does GPT-Image-2 support?
GPT-Image-2 supports text and image inputs and produces image outputs. It does not support audio or video input or output, and it does not natively provide text reasoning, coding, structured machine-readable output, function calling, or streaming.


Sources 6
Provider

About OpenAI