What is GPT-Image-2?
GPT-Image-2 is OpenAI’s image generation and editing model. It accepts text prompts and image inputs, then produces generated or edited images. The canonical API model identifier is gpt-image-2, and OpenAI also provides the dated snapshot gpt-image-2-2026-04-21 for applications that need a fixed model version.
The model is designed for workflows where the visual result matters more than producing a written answer. For example, a developer can use it to create a marketing illustration from a description, revise a product image using a reference picture, or generate an editorial graphic containing readable text. GPT-Image-2 is available through OpenAI’s Images API and also supports Batch API processing.
Where GPT-Image-2 fits in OpenAI’s lineup
GPT-Image-2 belongs to OpenAI’s GPT Image family and is focused specifically on visual generation and editing. It should not be treated as a smaller version of a general-purpose GPT model. Its primary output is an image, and the supplied specifications do not describe it as a text reasoning, coding, speech, or video model.
OpenAI’s current catalog also includes newer GPT Image 2.5 models for some workflows. The available research does not provide a complete feature or price comparison for those models, so GPT-Image-2 should be evaluated on its documented capabilities rather than assumed to be superior or inferior in every situation. Readers comparing the product family can also review GPT-Image-2.5 Sunburst and GPT-Image-2.5 Flare where those alternatives are relevant.
Supported inputs and outputs
GPT-Image-2 supports text input and image input. Image input is useful for editing, restyling, preserving reference details, or building a new composition around an existing visual. The model’s output is an image rather than a text response.
- Text input: Supported for describing the desired image or edit.
- Image input: Supported for reference-based generation and editing.
- Image output: Supported as the model’s primary result.
- Audio and video input: Not supported.
- Audio and video output: Not supported.
- Text, embedding, and structured machine-readable output: Not provided as native model outputs.
Image inputs are always processed at high fidelity. As a result, applications using GPT-Image-2 should omit the input_fidelity parameter rather than trying to select a lower or higher fidelity mode.
Resolution, quality and output formats
The model supports automatic sizing as well as custom resolutions. Custom dimensions must be multiples of 16. The longest edge cannot exceed 3,840 pixels, the aspect ratio cannot exceed 3:1, and the image must remain within OpenAI’s documented total-pixel limits. Resolutions above 2K are considered experimental, meaning that results can be more variable at larger sizes.
Generation quality can be set to low, medium, high, or auto. This gives developers a practical quality-versus-cost and quality-versus-speed control, although the supplied documentation does not provide a fixed generation time for each setting. Higher requested quality and larger images can also affect token consumption.
GPT-Image-2 supports PNG, JPEG, and WebP output. Transparent backgrounds are available in preview when using PNG or WebP. These format and transparency options are useful for product cutouts, interface assets, logos, stickers, composited marketing graphics, and other images that may need to be placed over a separate background.
GPT-Image-2 pricing and Batch API support
GPT-Image-2 uses token-based pricing rather than a single fixed price per generated image. The amount consumed depends on factors such as the prompt, reference-image inputs, requested resolution, and quality setting.
| Usage type | Published price |
|---|---|
| Image input | $8 per 1 million tokens |
| Cached image input | $2 per 1 million tokens |
| Image output | $30 per 1 million tokens |
| Text input | $5 per 1 million tokens |
| Cached text input | $1.25 per 1 million tokens |
| Text output, where applicable | $10 per 1 million tokens |
These rates do not translate directly into a universal per-image price because image dimensions, quality, prompts, and reference images change token usage. Developers should estimate costs using the actual workloads they expect to send rather than assuming every image has the same price.
Batch API requests receive a 50% discount according to OpenAI’s release documentation. Batch processing is therefore more appropriate for queued or non-interactive workloads, such as preparing a large set of catalog images or editorial assets, than for an application that must return an image immediately.
Main strengths and limitations
Strengths
- Reference-image handling: High-fidelity image input processing helps workflows that need to preserve details from supplied images.
- Text inside images: OpenAI describes improved text rendering, which is relevant to posters, labels, diagrams, advertisements, and other text-heavy visual assets.
- Flexible layouts: Custom resolutions and a broad aspect-ratio limit support outputs beyond standard square, portrait, and landscape presets.
- Production formats: PNG, JPEG, and WebP make the output practical for different delivery and editing requirements.
- Transparent backgrounds: Preview support for PNG and WebP can simplify compositing and asset preparation.
- Batch workflows: Batch API support and its discount can reduce the cost of large, asynchronous jobs.
Limitations
- GPT-Image-2 is not a general-purpose conversational model and does not provide native text reasoning or coding output.
- It does not support audio or video input or output.
- Function calling, structured outputs, and streaming are not supported according to the supplied model specifications.
- There is no documented context-window or maximum-output-token value for this image model.
- Pricing varies with token usage instead of following a simple fixed price per image.
- Resolutions above 2K are experimental, and custom image dimensions must follow strict multiple-of-16, pixel, edge, and aspect-ratio constraints.
Reasoning, coding and tool support
GPT-Image-2 should not be selected for tasks that require a reasoning model to analyze a problem and return a detailed text answer. The supplied specifications do not assign it a reasoning score or coding capability, and its text output is not a native model output. It is also not documented as supporting function calling, web search, or other tool use.
An application can still combine GPT-Image-2 with a separate language model or application layer. For example, another component could interpret a user’s request and produce an image prompt, while GPT-Image-2 handles the visual generation. That architecture does not make GPT-Image-2 itself a text reasoning or tool-using model; the responsibilities belong to different components.
Speed and cost trade-offs
The available editorial assessment rates GPT-Image-2’s speed at 8 out of 10 and cost at 6 out of 10. These are comparative editorial scores, not OpenAI-published benchmarks or guarantees. The documented controls provide a more concrete way to understand the trade-off: developers can choose low, medium, high, or automatic quality, while output size and image complexity influence token consumption.
For interactive applications, lower quality and moderate resolutions may be more appropriate when users need quick previews. High-quality, large-format images are better suited to final assets but may consume more tokens and produce less predictable latency. Batch processing is a separate option for workloads where immediate responses are unnecessary and the 50% discount is more valuable than real-time delivery.
Best use cases for GPT-Image-2
GPT-Image-2 is a strong fit for developers building image-focused products and workflows such as:
- Text-to-image creation tools for marketing, editorial, or creative teams.
- Reference-based product visualization and image revision.
- Advertising assets that require readable text within the image.
- Flexible-format social, web, packaging, or campaign graphics.
- Transparent product cutouts, stickers, icons, and composited design elements.
- Batch generation of catalog, campaign, or publishing assets.
- Applications where preserving important details from a supplied image is more important than generating an unrelated visual.
When to choose GPT-Image-2
Choose GPT-Image-2 when the central requirement is high-quality image generation or editing, especially when the workflow needs reference images, custom dimensions, readable embedded text, multiple output formats, or transparent backgrounds. It is also a reasonable choice when an application can manage token-based billing and does not need a fixed per-image price.
Consider another type of model when the main task is conversation, document analysis, coding, speech, video creation, structured JSON generation, or tool orchestration. A general-purpose text model is more appropriate for reasoning and code, while a video or audio model is needed for those media types. If the application requires streaming responses, function calling, or deterministic structured output, GPT-Image-2 is not the right standalone component.
GPT-Image-2 is therefore best understood as a specialized visual engine. Its advantage is control over image creation and editing rather than breadth across every modality. The most suitable alternative depends on whether the priority is lower cost, simpler fixed-image pricing, faster previews, richer text reasoning, or support for a different media format.

