What is GPT-5.4?
GPT-5.4 is OpenAI's flagship model for demanding professional work. It is designed to combine reasoning, coding, visual understanding, long-context processing, and tool use in one general-purpose model. The API model identifier is gpt-5.4, and OpenAI released it on March 5, 2026.
Unlike a model intended only for short conversational answers, GPT-5.4 is aimed at multi-step tasks such as software development, research synthesis, spreadsheet and presentation work, document analysis, browser-based operations, and workflows that require an AI system to use external tools. It is available through the OpenAI API, ChatGPT, and Codex.
Its position in OpenAI's current lineup is that of a high-capability model for complex work. That positioning brings a larger context window and stronger reasoning-oriented features, but it also means that GPT-5.4 is not automatically the best choice for every request. Simpler, high-volume tasks may be better served by a smaller or less expensive model.
Inputs, outputs, and supported capabilities
GPT-5.4 accepts text and image inputs and produces text output. Image input allows the model to interpret visual material such as documents, charts, screenshots, and other supported images. The model does not natively produce images, audio, or video as its primary output.
OpenAI may expose image generation, code execution, web search, file search, hosted shell, and computer use as tools in compatible workflows. These tools extend what an application can accomplish, but they should not be confused with native non-text output from GPT-5.4 itself. For example, an image-generation tool may create an image in a workflow, while the GPT-5.4 model remains a text-output model.
- Text input and text output
- Image input with high-fidelity visual understanding
- Configurable reasoning effort
- Function calling and structured outputs
- Web search, file search, code interpreter, and hosted shell tools
- Computer-use actions involving screenshots, mouse input, and keyboard input
- MCP and dynamic tool search
- Streaming responses and batch processing
The model supports the Responses and Chat Completions APIs. Its tool support is particularly relevant to applications that need more than generated text, such as research agents, coding assistants, document workflows, and software agents operating in controlled environments.
Context window and output limits
GPT-5.4 has a context window of 1,050,000 tokens and supports a maximum output of 128,000 tokens. A context window is the amount of information the model can consider in a request and its surrounding conversation, including supplied documents, tool results, and other input.
The unusually large context is useful for long codebases, substantial collections of documents, extended research sessions, and multi-step agent workflows. It does not mean that every long prompt will be equally effective: applications still need to organize information carefully, remove irrelevant material, and control tool results to avoid unnecessary cost and complexity.
OpenAI documents a 272,000-token threshold for extended-context pricing. Requests that exceed that threshold are charged at twice the standard input rate and 1.5 times the standard output rate for the full session under standard, batch, and Flex processing. This pricing rule is separate from the model's maximum context capacity.
Reasoning and coding performance
GPT-5.4 supports five reasoning effort settings: none, low, medium, high, and xhigh. Reasoning effort controls how much additional internal work the model applies to a request. Lower settings can be appropriate for straightforward tasks, while higher settings are intended for difficult analysis, planning, debugging, and multi-step problem solving.
Higher reasoning settings may increase latency and token usage. The practical choice is therefore not simply to select the highest setting for every request. A production application might use a lower setting for routine transformations and reserve high or xhigh reasoning for tasks where improved analysis justifies additional time and cost.
GPT-5.4 also incorporates coding advances associated with GPT-5.3-Codex while extending them to broader professional and computer-use workflows. It can assist with software engineering, code review, debugging, repository analysis, and multi-step implementation tasks. Its coding value is greatest when the application supplies relevant files, test results, documentation, or tools that let the model inspect and act on the development environment.
Computer use, web search, and tool calling
Function calling allows an application to give the model defined operations that it can request during a response. Structured outputs can help applications receive data in a predictable format. Together, these capabilities support workflows where GPT-5.4 interprets a request, decides which operation is needed, and incorporates the result into its answer or next action.
Computer-use support allows compatible agents to operate software environments through screenshots and mouse and keyboard actions. This can support browser-based tasks, desktop workflows, and applications that do not expose every operation through a conventional API. Computer use should be deployed with appropriate permissions, confirmation steps, and monitoring because an agent's actions can affect external systems.
GPT-5.4 also supports web search and other hosted tools. Web search is important when a task depends on information newer than the model's knowledge cutoff, which is August 31, 2025. Retrieval tools can provide current information during use, but they do not change the underlying cutoff of the model.
Dynamic tool search can load relevant tools when they are needed instead of placing every available tool definition in the initial context. This can reduce the amount of tool-description material competing for context and may be useful in applications with many connected tools or MCP servers.
GPT-5.4 API pricing
OpenAI's standard API pricing is $2.50 per 1 million input tokens, $0.25 per 1 million cached input tokens, and $15.00 per 1 million output tokens.
| Usage type | Standard price |
|---|---|
| Input tokens | $2.50 per 1 million tokens |
| Cached input tokens | $0.25 per 1 million tokens |
| Output tokens | $15.00 per 1 million tokens |
Cached input pricing applies when eligible prompt content can be reused according to OpenAI's prompt-caching rules. Batch and Flex processing are available at half the standard rate, while priority processing costs twice the standard rate. Regional processing endpoints carry a 10% uplift.
Applications should budget separately for input, cached input, output, tool activity, and any extended-context surcharge. The output rate is substantially higher than the ordinary input rate, so concise prompts and controlled response lengths can matter for workloads that generate large amounts of text.
Main strengths and limitations
GPT-5.4's main strength is the combination of capabilities in one model. A single workflow can use text reasoning, image understanding, long-context input, code generation, web retrieval, structured tool calls, and computer-use actions. This reduces the need to divide a complex task among several specialized components.
The model is particularly well suited to long-horizon work: tasks in which the system must understand substantial context, make a plan, use tools, inspect results, and continue through several stages. Its large context window can also reduce the need to summarize or split large source collections before analysis.
There are important limitations. GPT-5.4's knowledge cutoff is August 31, 2025, so current facts should be retrieved rather than assumed. Outputs can still be incorrect or overconfident, especially when a task involves ambiguous source material or consequential decisions. Tool access improves an application's ability to retrieve information or take action, but it does not remove the need for validation and access controls.
GPT-5.4 does not support fine-tuning according to the supplied model documentation. It also does not natively generate images, audio, or video. Although it supports visual input and can use compatible tools, those capabilities should not be described as native image, audio, or video output.
Best use cases for GPT-5.4
- Complex software engineering: Repository analysis, implementation planning, debugging, code review, and tasks that require repeated interaction with development tools.
- Long-document and visual analysis: Reviewing large document sets, extracting information from visual documents, and combining text with charts or screenshots.
- Research and synthesis: Comparing sources, organizing evidence, and using web search when information must be current.
- Agentic workflows: Multi-step processes that involve planning, tool selection, function calls, and controlled computer interaction.
- Business productivity: Spreadsheet, presentation, and document workflows where the model must reason about several related files or actions.
- Structured application output: Producing predictable fields for downstream software through structured outputs and function calling.
When should you choose GPT-5.4?
Choose GPT-5.4 when the task benefits from advanced reasoning, a very large context window, image understanding, coding, or multiple tools in the same workflow. It is a strong candidate for an agent that must work through a lengthy task rather than answer a single simple question.
Its capability comes with trade-offs. GPT-5.4 is less appropriate for simple classification, basic extraction, very low-latency responses, or the lowest possible inference cost. For those workloads, a smaller model may provide sufficient quality with faster responses and lower per-token spending. The supplied research does not specify particular smaller GPT-5.4 sibling prices or limits, so the choice should be made using the current documentation and measured workload costs rather than assumed family-wide specifications.
Higher reasoning settings should also be used selectively. A request that only needs a short transformation may not benefit from high or xhigh reasoning. Conversely, difficult planning, debugging, and tool-heavy tasks may justify the added latency and token consumption.
Practical evaluation checklist
Before deploying GPT-5.4, test it on representative tasks rather than relying only on general capability descriptions. Measure answer quality, tool-selection accuracy, completion time, input and output token use, and the frequency of required human review.
For long-context applications, test whether the model can locate and use relevant information across the full document set. For computer-use workflows, evaluate failure recovery and enforce limits on what the agent can click, edit, purchase, send, or execute. For current-information tasks, verify web-search citations or retrieved evidence instead of treating the model's stored knowledge as current.
GPT-5.4 is best understood as a high-capability option for complex, tool-connected work. Its large context, reasoning controls, coding support, vision input, and agent features can justify its cost when they replace manual coordination or several separate model steps. For routine or price-sensitive workloads, a faster and less expensive option may be the better engineering decision.

