What is GPT-Image-1?
GPT-Image-1 is OpenAI’s image-generation and image-editing model for API use. It accepts text prompts and image inputs, then produces image outputs. In practical terms, a developer can use it to create an illustration from a description, revise an existing image, preserve visual elements from reference images, or replace a selected area through a mask-based editing workflow.
OpenAI released GPT-Image-1 on April 23, 2025. The company described it as the API model behind the native image-generation experience introduced in ChatGPT. Its intended role is visual content production rather than general conversation, reasoning, coding, speech, video generation, or embeddings.
The model is currently listed as deprecated. OpenAI’s model documentation states that API access is scheduled to end on October 23, 2026. That lifecycle status is one of the most important practical considerations: GPT-Image-1 can still be relevant for understanding or maintaining an existing integration, but it is a less suitable foundation for a new long-lived product unless a migration path has been considered.
Core image-generation and editing capabilities
GPT-Image-1 supports both text-to-image generation and image editing. A text-only request can describe the desired subject, composition, style, or design. When an image is supplied, it can act as a reference for the generated result or as the source for an edit. Masks allow an application to identify the portion of an image that should be changed, which is useful for inpainting-style workflows such as replacing an object, altering a background, or correcting a localized area.
- Text-to-image generation: Create a new image from a written prompt.
- Reference-image workflows: Use supplied images to guide appearance, composition, or identity.
- Mask-based editing: Target particular regions instead of changing the entire image.
- Transparent backgrounds: Generate assets with transparency when a transparent background is requested and PNG or WebP output is selected.
- Quality controls: Select low, medium, high, or automatic quality.
- Output formats: Receive PNG, JPEG, or WebP images, with optional JPEG and WebP compression.
These features make the model more useful for production workflows than a generator that only creates a new image from a prompt. For example, an e-commerce application could create product artwork, revise a supplied product image, or generate a transparent asset for placement in a catalog.
Supported sizes, quality, and formats
The documented image sizes are 1024×1024, 1024×1536, and 1536×1024. An automatic size option is also available. The square setting is suitable for many social, thumbnail, and product-card uses, while the portrait and landscape settings provide more room for posters, banners, and editorial compositions.
Quality can be set to low, medium, high, or auto. Higher quality generally carries a higher image-generation charge, so the setting should match the task. Low quality can be appropriate for drafts or rapid iteration, whereas high quality may be justified for a final marketing asset. The provider documents the following per-image generation prices:
| Quality | 1024×1024 | 1024×1536 or 1536×1024 |
|---|---|---|
| Low | $0.011 per image | $0.016 per image |
| Medium | $0.042 per image | $0.063 per image |
| High | $0.167 per image | $0.25 per image |
PNG, JPEG, and WebP are supported. JPEG and WebP compression can be configured where appropriate, while transparent backgrounds require an explicit transparent-background request together with PNG or WebP output.
GPT-Image-1 pricing and cost planning
GPT-Image-1 pricing has more than one component. The per-image prices above cover the image-generation charge, but a request can also incur charges for text tokens and image tokens. OpenAI documents these rates:
- Text input: $5 per million tokens.
- Cached text input: $1.25 per million tokens.
- Image input: $10 per million image tokens.
- Cached image input: $2.50 per million image tokens.
- Image output: $40 per million image tokens.
This distinction matters most for editing and reference-image workflows. A request that supplies one or more images can incur image-input charges in addition to the selected generation price. Applications that repeatedly reuse the same inputs may benefit from the documented cached-input rates, but the final cost still depends on the request’s token usage and image settings.
From a cost perspective, low-quality square images are the least expensive documented generation option. High-quality portrait or landscape images are the most expensive per-image option. These are provider-published prices, not an estimate of total application cost; storage, retries, orchestration, and any other application expenses are separate.
Inputs, outputs, and unsupported features
GPT-Image-1 has multimodal input because it can process both text and images. Its direct output is image data, not text, audio, or video. The supplied model documentation does not publish a context-length limit or maximum-output-token limit for this model; those values should therefore be treated as undocumented rather than assumed to be unlimited.
GPT-Image-1 is not documented as a general-purpose language model. It does not support text generation as its primary output, and it should not be selected for ordinary writing, coding, long-form reasoning, speech, video, or embedding workloads. The model documentation also lists the following unsupported features:
- Streaming
- Function calling
- Structured outputs
- Fine-tuning
There is consequently no documented reasoning score or coding capability to evaluate in the same way as a text-generation model. Any reasoning involved in interpreting an image prompt should not be confused with a supported general reasoning interface.
Strengths and limitations
Where GPT-Image-1 is strongest
GPT-Image-1’s main strength is the combination of generation and editing in one API-oriented image model. It can create new visual assets, use reference images, perform localized mask-based changes, and return common web-friendly formats. The documented quality and size controls also let developers trade visual quality, dimensions, and cost for different stages of a workflow.
It is particularly well suited to marketing graphics, e-commerce imagery, illustrations, concept art, educational visuals, gaming assets, and applications that let users revise their own images. Transparent-background support is useful for logos, product cutouts, icons, characters, and other assets that need to be placed on different backgrounds.
Important limitations
The largest limitation is its lifecycle. GPT-Image-1 is deprecated and scheduled to shut down on October 23, 2026. A team choosing it for a new application should plan for replacement rather than treating the model as a permanent endpoint.
It also lacks several capabilities that may matter in a broader automation system. There is no streaming, function calling, structured-output interface, or fine-tuning support documented for this exact model. It is not a text, speech, video, or embedding model, so a product that needs those functions will require additional models or services. Editing and reference-image requests can also cost more than the headline per-image price because image tokens are billed separately.
When to choose GPT-Image-1
GPT-Image-1 is a reasonable choice when the central requirement is API-based image creation or editing and the project can accommodate its announced shutdown date. It is especially appropriate when the workflow needs one or more of the following:
- Generating images from natural-language descriptions.
- Editing user-provided images with references or masks.
- Creating marketing, product, educational, or game-related visuals.
- Producing transparent-background PNG or WebP assets.
- Choosing between low, medium, and high generation quality.
- Using square, portrait, or landscape output dimensions.
It is less appropriate when the application needs a general conversational model, coding or reasoning, real-time streaming, function calls, structured JSON responses, audio or video, or a stable long-term model without migration work. In those cases, a current model designed for the required modality or tool interface is more suitable. The supplied research does not identify a specific successor by name, so comparisons should be made against the current GPT Image offering documented by OpenAI rather than assuming an unverified replacement.
Practical evaluation before adoption
Before integrating GPT-Image-1, test the exact prompts and image-editing cases that matter to the product. Compare low, medium, and high quality at the required dimensions, and measure the resulting per-request cost rather than relying only on the generation fee. Include reference-image and mask-based requests in testing because their image-token usage can change the economics.
Also verify that the application can handle image responses in the selected format, transparent-background requirements, compression settings, and the absence of streaming or structured outputs. Most importantly, document a migration plan before launch. Because OpenAI has announced a shutdown date, a successful short-term integration still needs an exit strategy.

