What is GPT-4o?
GPT-4o is OpenAI’s general-purpose model for text generation and image understanding. The “o” stands for “omni,” reflecting OpenAI’s positioning of the model as a more broadly capable GPT system. OpenAI introduced GPT-4o on May 13, 2024, describing it as a flagship model with GPT-4-level intelligence, improved multilingual performance, stronger vision capabilities, and lower API pricing than GPT-4 Turbo.
For current API use, the canonical gpt-4o model accepts text and image inputs and produces text. That text can be ordinary natural-language output or structured output that follows a supported JSON schema. GPT-4o’s ability to interpret images makes it suitable for tasks such as analyzing documents, answering questions about charts, extracting information from screenshots, and combining written instructions with visual material.
GPT-4o is not an image-generation model, and the current standard model page does not identify it as an audio or video input/output model. Image generation, speech, transcription, and video generation require separate models or services where available.
GPT-4o specifications at a glance
| Specification | GPT-4o |
|---|---|
| Provider | OpenAI |
| Release date | May 13, 2024 |
| Model identifier | gpt-4o |
| Context window | 128,000 tokens |
| Maximum output | 16,384 tokens |
| Knowledge cutoff | October 1, 2023 |
| Inputs | Text and images |
| Output | Text |
| Standard input price | $2.50 per 1 million tokens |
| Cached input price | $1.25 per 1 million tokens |
| Standard output price | $10 per 1 million tokens |
The context window is the amount of input and generated content the model can handle in a request, subject to the limits of the specific endpoint and application. The maximum output limit is separate: a request may include a large amount of context, but GPT-4o can generate no more than 16,384 output tokens for the supported configuration.
Capabilities and supported modalities
GPT-4o’s main API capability is text generation from text or image input. This supports ordinary conversation, summarization, classification, extraction, rewriting, coding assistance, multilingual generation, and visual question answering. For example, an application can provide a product photograph and ask for a description, submit a chart for interpretation, or pass a scanned document for structured field extraction.
Image understanding is not image generation
GPT-4o can analyze images, but it does not directly create images. A workflow that needs both visual analysis and image creation would typically use GPT-4o for interpretation and a separate OpenAI image model or image-generation endpoint for the output image. The same distinction applies to audio and video: the standard GPT-4o model should not be treated as a native speech, audio, or video generation system.
Knowledge cutoff and current information
OpenAI identifies GPT-4o’s knowledge cutoff as October 1, 2023. This is the underlying training cutoff and does not change when an application supplies additional context. The base model also should not be treated as a current web-search model. If an application needs up-to-date information, it should provide retrieved material or use a separately supported search workflow rather than assume that GPT-4o’s built-in knowledge is current.
Developer features and tool support
GPT-4o supports function calling, streaming, structured outputs, and batch processing. Function calling lets an application describe external actions—such as looking up an account, querying a database, or creating a task—and allows the model to request one of those actions in a defined format. The application remains responsible for executing the function and validating its arguments.
Streaming sends generated output incrementally instead of waiting for the entire response. This can make an assistant feel more responsive, particularly when the model is producing a long answer. Batch processing is intended for workloads that can be handled asynchronously rather than requiring an immediate interactive response.
Structured Outputs can constrain a response to a supported JSON schema, which is useful for reliable extraction and downstream software processing. This is different from JSON mode. JSON mode can help return valid JSON, but it does not by itself guarantee that the response follows a particular schema. Applications should still validate model output and handle refusals, missing fields, and unexpected values.
OpenAI also lists fine-tuning as supported for GPT-4o. However, the fine-tuning platform is being wound down: new users no longer have access, while some existing users may retain limited training access during the transition. Existing fine-tuned models remain available for inference until their underlying base models are deprecated. This makes fine-tuning a lifecycle consideration rather than an assumption that every new project can start training a custom GPT-4o model.
GPT-4o pricing and cost trade-offs
Current standard API pricing is $2.50 per 1 million input tokens and $10 per 1 million output tokens. Cached input tokens are priced at $1.25 per 1 million tokens. These rates make input reuse less expensive when an application repeatedly sends eligible prompt content, although actual charges depend on the endpoint, token usage, caching conditions, and workload.
Output is four times more expensive per token than standard input at the listed rates. Applications can therefore reduce cost by avoiding unnecessarily long responses, limiting output length where appropriate, reusing stable prompt content through caching, and selecting a less capable or lower-cost model when the task does not need GPT-4o’s image understanding or general capability level.
GPT-4o’s practical trade-off is capability versus specialized optimization. It is broader than a narrowly focused extraction or classification system, but a dedicated smaller model may be more economical for repetitive, simple tasks. Conversely, a newer or specialized reasoning model may be more appropriate for difficult multi-step reasoning even if it has different speed or cost characteristics. The supplied research does not establish a direct price or benchmark comparison with those alternatives, so such comparisons should be checked against current provider documentation.
Reasoning, coding, and performance
GPT-4o is designed as a general-purpose model rather than a dedicated extended-reasoning model. It can explain problems, follow multi-step instructions, generate code, review code, and assist with debugging, but the available research does not provide a standardized benchmark proving that it is the strongest option for every reasoning task. For complex problems where deliberate reasoning quality is more important than response speed, a dedicated reasoning model may be a better fit.
For coding, GPT-4o is useful for generating functions, explaining unfamiliar code, transforming data formats, writing tests, and helping investigate errors. Function calling and structured outputs also make it suitable for software workflows in which the model must return predictable arguments for an application to validate. Generated code should still be reviewed and tested, especially when it changes data, accesses external systems, or handles security-sensitive information.
OpenAI’s model information characterizes GPT-4o as relatively fast. The supplied editorial assessment rates its speed highly and its reasoning and coding capabilities as strong, but those are evaluations rather than provider-published scores. Performance will vary with prompt design, input size, image complexity, output length, endpoint, and workload.
Availability and model lifecycle
GPT-4o was retired from ChatGPT on February 13, 2026. That retirement applies to the consumer ChatGPT product and does not, according to the supplied research, mean that the canonical API model has been removed. The gpt-4o alias remains listed as an OpenAI API model.
Developers should distinguish the alias from dated snapshots. The gpt-4o-2024-05-13 snapshot is scheduled for API shutdown on October 23, 2026. Other dated identities, including gpt-4o-2024-08-06 and gpt-4o-2024-11-20, are separate snapshots rather than interchangeable names. A production integration should record the exact model ID, endpoint, and lifecycle information instead of assuming that ChatGPT availability and API availability are identical.
The model also should not be confused with related specialized products such as GPT-4o Search Preview, GPT-4o Realtime, or GPT-4o Audio. Those names describe different model or product identities with different capabilities and lifecycle conditions.
Best use cases for GPT-4o
GPT-4o is a practical choice when one API model needs to handle several types of general-purpose work:
- Image-aware assistants: Answer questions about photographs, screenshots, diagrams, charts, and other supported visual inputs.
- Document processing: Extract fields, classify documents, summarize content, or turn visual and textual material into structured records.
- Business automation: Combine natural-language requests with function calls to internal systems and return predictable structured results.
- Coding assistance: Generate examples, explain code, suggest fixes, and help create tests or data transformations.
- Multilingual applications: Produce and transform text across languages where the application needs a broad general-purpose model.
- Latency-sensitive workflows: Use a model with broad capabilities without automatically choosing a slower dedicated reasoning system for every request.
When to choose GPT-4o
Choose GPT-4o when your application needs text generation plus image understanding, structured extraction, function calling, or a combination of these features in one relatively fast API model. It is especially suitable when the task is complex enough to benefit from a GPT-4-class model but does not require the strongest available specialized reasoning system.
Consider another option when the primary requirement is different:
- Use a dedicated reasoning model when difficult multi-step reasoning is more important than general-purpose speed and modality support.
- Use a separate image model when the application must generate or edit images.
- Use dedicated audio, speech, or video systems when those modalities are central to input or output.
- Use a current retrieval or web-search workflow when answers depend on information newer than the October 1, 2023 knowledge cutoff.
- Use a lower-cost or smaller model for simple, high-volume tasks that do not need GPT-4o’s image understanding or broader capabilities.
- Use a newer model such as GPT-4.1 when its documented capabilities and lifecycle better match the application’s requirements.
GPT-4o’s main advantage is breadth: it combines text generation, image input, developer controls, and relatively responsive performance. Its main limitations are equally important: text-only output, no current-information guarantee without retrieval, a finite knowledge cutoff, lifecycle differences between ChatGPT and the API, and the need to verify exact model and snapshot status before deployment.

