GPT-5.4

GPT-5.4

by OpenAI · Current

GPT-5.4 is OpenAI's flagship professional-work model with reasoning, coding, image understanding, computer use, web search, structured outputs, and a 1.05 million-token context window. It costs $2.50 per million input tokens and $15 per million output tokens.

Text Actions Reasoning Coding
Released on March 5, 2026, GPT-5.4 is OpenAI's general-purpose frontier model for professional workflows, coding, research, document-heavy tasks, and computer-using agents. It accepts text and image inputs, produces text outputs, supports up to 1.05 million tokens of context and 128,000 output tokens, and is available through the Responses and Chat Completions APIs.
Outputs

What GPT-5.4 can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

10/10 Reasoning
9/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family GPT-5.4
Model type Reasoning
Context window 1.05M tokens
Maximum output 128K tokens
Knowledge cutoff August 31, 2025
Release date March 5, 2026
Status Current
Knowledge cutoff notes

The official GPT-5.4 model documentation lists August 31, 2025 as the model's knowledge cutoff. Web search and other retrieval tools can provide newer information during use but do not change the underlying cutoff.

Model notes

GPT-5.4 is available through the OpenAI API as gpt-5.4 and has a dated snapshot named gpt-5.4-2026-03-05. It supports reasoning effort values none, low, medium, high, and xhigh. The model accepts text and image input and produces text output. Image generation, code interpreter, hosted shell, computer use, web search, file search, MCP, tool search, and other capabilities are exposed as supported tools rather than native non-text model outputs. Standard pricing is $2.50 per 1 million input tokens, $0.25 per 1 million cached input tokens, and $15 per 1 million output tokens. Batch and Flex processing are available at half the standard rate, while priority processing is twice the standard rate. Prompts exceeding 272,000 input tokens are subject to extended-context pricing of 2x input and 1.5x output for the full session. Regional processing endpoints have a 10% uplift. Editorial scores are comparative estimates, not OpenAI specifications.

Cost

Model pricing

Input $2.50 per 1 million input tokens; $0.25 per 1 million cached input tokens
Output $15.00 per 1 million output tokens
Model guide

GPT-5.4: Features, Pricing, Context Window, and API Capabilities

GPT-5.4 is OpenAI's flagship reasoning model for complex professional work, combining advanced reasoning, coding, multimodal vision, long-context processing, computer use, web search, and agentic tool calling in a single API model.

What is GPT-5.4?

GPT-5.4 is OpenAI's flagship model for demanding professional work. It is designed to combine reasoning, coding, visual understanding, long-context processing, and tool use in one general-purpose model. The API model identifier is gpt-5.4, and OpenAI released it on March 5, 2026.

Unlike a model intended only for short conversational answers, GPT-5.4 is aimed at multi-step tasks such as software development, research synthesis, spreadsheet and presentation work, document analysis, browser-based operations, and workflows that require an AI system to use external tools. It is available through the OpenAI API, ChatGPT, and Codex.

Its position in OpenAI's current lineup is that of a high-capability model for complex work. That positioning brings a larger context window and stronger reasoning-oriented features, but it also means that GPT-5.4 is not automatically the best choice for every request. Simpler, high-volume tasks may be better served by a smaller or less expensive model.

Inputs, outputs, and supported capabilities

GPT-5.4 accepts text and image inputs and produces text output. Image input allows the model to interpret visual material such as documents, charts, screenshots, and other supported images. The model does not natively produce images, audio, or video as its primary output.

OpenAI may expose image generation, code execution, web search, file search, hosted shell, and computer use as tools in compatible workflows. These tools extend what an application can accomplish, but they should not be confused with native non-text output from GPT-5.4 itself. For example, an image-generation tool may create an image in a workflow, while the GPT-5.4 model remains a text-output model.

  • Text input and text output
  • Image input with high-fidelity visual understanding
  • Configurable reasoning effort
  • Function calling and structured outputs
  • Web search, file search, code interpreter, and hosted shell tools
  • Computer-use actions involving screenshots, mouse input, and keyboard input
  • MCP and dynamic tool search
  • Streaming responses and batch processing

The model supports the Responses and Chat Completions APIs. Its tool support is particularly relevant to applications that need more than generated text, such as research agents, coding assistants, document workflows, and software agents operating in controlled environments.

Context window and output limits

GPT-5.4 has a context window of 1,050,000 tokens and supports a maximum output of 128,000 tokens. A context window is the amount of information the model can consider in a request and its surrounding conversation, including supplied documents, tool results, and other input.

The unusually large context is useful for long codebases, substantial collections of documents, extended research sessions, and multi-step agent workflows. It does not mean that every long prompt will be equally effective: applications still need to organize information carefully, remove irrelevant material, and control tool results to avoid unnecessary cost and complexity.

OpenAI documents a 272,000-token threshold for extended-context pricing. Requests that exceed that threshold are charged at twice the standard input rate and 1.5 times the standard output rate for the full session under standard, batch, and Flex processing. This pricing rule is separate from the model's maximum context capacity.

Reasoning and coding performance

GPT-5.4 supports five reasoning effort settings: none, low, medium, high, and xhigh. Reasoning effort controls how much additional internal work the model applies to a request. Lower settings can be appropriate for straightforward tasks, while higher settings are intended for difficult analysis, planning, debugging, and multi-step problem solving.

Higher reasoning settings may increase latency and token usage. The practical choice is therefore not simply to select the highest setting for every request. A production application might use a lower setting for routine transformations and reserve high or xhigh reasoning for tasks where improved analysis justifies additional time and cost.

GPT-5.4 also incorporates coding advances associated with GPT-5.3-Codex while extending them to broader professional and computer-use workflows. It can assist with software engineering, code review, debugging, repository analysis, and multi-step implementation tasks. Its coding value is greatest when the application supplies relevant files, test results, documentation, or tools that let the model inspect and act on the development environment.

Computer use, web search, and tool calling

Function calling allows an application to give the model defined operations that it can request during a response. Structured outputs can help applications receive data in a predictable format. Together, these capabilities support workflows where GPT-5.4 interprets a request, decides which operation is needed, and incorporates the result into its answer or next action.

Computer-use support allows compatible agents to operate software environments through screenshots and mouse and keyboard actions. This can support browser-based tasks, desktop workflows, and applications that do not expose every operation through a conventional API. Computer use should be deployed with appropriate permissions, confirmation steps, and monitoring because an agent's actions can affect external systems.

GPT-5.4 also supports web search and other hosted tools. Web search is important when a task depends on information newer than the model's knowledge cutoff, which is August 31, 2025. Retrieval tools can provide current information during use, but they do not change the underlying cutoff of the model.

Dynamic tool search can load relevant tools when they are needed instead of placing every available tool definition in the initial context. This can reduce the amount of tool-description material competing for context and may be useful in applications with many connected tools or MCP servers.

GPT-5.4 API pricing

OpenAI's standard API pricing is $2.50 per 1 million input tokens, $0.25 per 1 million cached input tokens, and $15.00 per 1 million output tokens.

Usage typeStandard price
Input tokens$2.50 per 1 million tokens
Cached input tokens$0.25 per 1 million tokens
Output tokens$15.00 per 1 million tokens

Cached input pricing applies when eligible prompt content can be reused according to OpenAI's prompt-caching rules. Batch and Flex processing are available at half the standard rate, while priority processing costs twice the standard rate. Regional processing endpoints carry a 10% uplift.

Applications should budget separately for input, cached input, output, tool activity, and any extended-context surcharge. The output rate is substantially higher than the ordinary input rate, so concise prompts and controlled response lengths can matter for workloads that generate large amounts of text.

Main strengths and limitations

GPT-5.4's main strength is the combination of capabilities in one model. A single workflow can use text reasoning, image understanding, long-context input, code generation, web retrieval, structured tool calls, and computer-use actions. This reduces the need to divide a complex task among several specialized components.

The model is particularly well suited to long-horizon work: tasks in which the system must understand substantial context, make a plan, use tools, inspect results, and continue through several stages. Its large context window can also reduce the need to summarize or split large source collections before analysis.

There are important limitations. GPT-5.4's knowledge cutoff is August 31, 2025, so current facts should be retrieved rather than assumed. Outputs can still be incorrect or overconfident, especially when a task involves ambiguous source material or consequential decisions. Tool access improves an application's ability to retrieve information or take action, but it does not remove the need for validation and access controls.

GPT-5.4 does not support fine-tuning according to the supplied model documentation. It also does not natively generate images, audio, or video. Although it supports visual input and can use compatible tools, those capabilities should not be described as native image, audio, or video output.

Best use cases for GPT-5.4

  • Complex software engineering: Repository analysis, implementation planning, debugging, code review, and tasks that require repeated interaction with development tools.
  • Long-document and visual analysis: Reviewing large document sets, extracting information from visual documents, and combining text with charts or screenshots.
  • Research and synthesis: Comparing sources, organizing evidence, and using web search when information must be current.
  • Agentic workflows: Multi-step processes that involve planning, tool selection, function calls, and controlled computer interaction.
  • Business productivity: Spreadsheet, presentation, and document workflows where the model must reason about several related files or actions.
  • Structured application output: Producing predictable fields for downstream software through structured outputs and function calling.

When should you choose GPT-5.4?

Choose GPT-5.4 when the task benefits from advanced reasoning, a very large context window, image understanding, coding, or multiple tools in the same workflow. It is a strong candidate for an agent that must work through a lengthy task rather than answer a single simple question.

Its capability comes with trade-offs. GPT-5.4 is less appropriate for simple classification, basic extraction, very low-latency responses, or the lowest possible inference cost. For those workloads, a smaller model may provide sufficient quality with faster responses and lower per-token spending. The supplied research does not specify particular smaller GPT-5.4 sibling prices or limits, so the choice should be made using the current documentation and measured workload costs rather than assumed family-wide specifications.

Higher reasoning settings should also be used selectively. A request that only needs a short transformation may not benefit from high or xhigh reasoning. Conversely, difficult planning, debugging, and tool-heavy tasks may justify the added latency and token consumption.

Practical evaluation checklist

Before deploying GPT-5.4, test it on representative tasks rather than relying only on general capability descriptions. Measure answer quality, tool-selection accuracy, completion time, input and output token use, and the frequency of required human review.

For long-context applications, test whether the model can locate and use relevant information across the full document set. For computer-use workflows, evaluate failure recovery and enforce limits on what the agent can click, edit, purchase, send, or execute. For current-information tasks, verify web-search citations or retrieved evidence instead of treating the model's stored knowledge as current.

GPT-5.4 is best understood as a high-capability option for complex, tool-connected work. Its large context, reasoning controls, coding support, vision input, and agent features can justify its cost when they replace manual coordination or several separate model steps. For routine or price-sensitive workloads, a faster and less expensive option may be the better engineering decision.


Answers to Frequently Asked Questions

When should you choose GPT-5.4 instead of a smaller model?
Choose GPT-5.4 for complex, multi-step tasks that benefit from advanced reasoning, a very large context window, image understanding, coding, or multiple tools. A smaller model may be more suitable for simple classification, basic extraction, low-latency responses, or cost-sensitive high-volume workloads.
What reasoning and tool-use capabilities does GPT-5.4 support?
GPT-5.4 offers five reasoning settings—none, low, medium, high, and xhigh—and supports function calling, structured outputs, web search, file search, code interpreter, hosted shell, MCP, dynamic tool search, and computer-use actions involving screenshots, mouse input, and keyboard input.
How much does GPT-5.4 cost through the API?
Standard GPT-5.4 API pricing is $2.50 per 1 million input tokens, $0.25 per 1 million cached input tokens, and $15.00 per 1 million output tokens. Batch and Flex processing cost half the standard rate, while priority processing costs twice the standard rate.
What is GPT-5.4 designed for?
GPT-5.4 is OpenAI's flagship model for demanding professional work, including advanced reasoning, software development, research synthesis, long-document analysis, spreadsheet and presentation workflows, visual understanding, and tool-connected agentic tasks.
What are GPT-5.4's context window and output limits?
GPT-5.4 has a context window of 1,050,000 tokens and supports a maximum output of 128,000 tokens. Requests exceeding 272,000 tokens are subject to extended-context pricing at higher rates.


Sources 5
Provider

About OpenAI