GPT-5.4

GPT-5.4 Mini

by OpenAI · Current

GPT-5.4 Mini is a fast, cost-efficient OpenAI reasoning model with a 400,000-token context window, 128,000-token maximum output, image input, structured outputs, streaming, web search, computer use, and broad Responses API tool support.

Text Reasoning Coding
GPT-5.4 Mini is a current OpenAI model released on March 17, 2026. It combines reasoning, coding, multimodal image understanding, and tool-use capabilities with lower latency and lower token prices than GPT-5.4, making it suitable for high-volume applications and responsive agent workflows.
Outputs

What GPT-5.4 Mini can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming JSON mode Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-5.4
Model type Lightweight
Context window 400K tokens
Maximum output 128K tokens
Knowledge cutoff 2025-08-31
Release date 2026-03-17
Status Current
Knowledge cutoff notes

The official model page lists August 31, 2025 as the model's knowledge cutoff. This cutoff is distinct from web search or other external tools that can provide newer information during use.

Model notes

GPT-5.4 Mini supports reasoning_effort values none, low, medium, high, and xhigh. The canonical API model ID is gpt-5.4-mini, with the dated snapshot gpt-5.4-mini-2026-03-17. It accepts text and image inputs and returns text. Audio and video inputs are not supported. Responses API tools include web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Regional processing endpoints have a 10% uplift. The editorial scores are comparative estimates rather than official OpenAI ratings.

Cost

Model pricing

Input $0.75 per 1 million input tokens; $0.075 per 1 million cached input tokens
Output $4.50 per 1 million output tokens
Model guide

GPT-5.4 Mini: Capabilities, Pricing, Context Window and API Support

GPT-5.4 Mini is OpenAI's faster, lower-cost GPT-5.4 model for coding, computer use, tool-enabled applications, subagents, and high-volume workloads. It supports text and image inputs, text output, reasoning effort controls, a 400,000-token context window, 128,000 maximum output tokens, structured outputs, streaming, web search, and multiple Responses API tools.

GPT-5.4 Mini is OpenAI's lightweight member of the GPT-5.4 model family. It is designed for applications that need useful reasoning and coding performance but cannot justify the latency or cost of a larger general-purpose model. OpenAI positions it for coding, computer-use agents, subagents, tool-enabled workflows, and other workloads that may generate many model requests.

The model was released on March 17, 2026, and is available through the OpenAI API. Its canonical model ID is gpt-5.4-mini; OpenAI also lists the dated snapshot gpt-5.4-mini-2026-03-17. The specifications below distinguish documented model properties from editorial assessments of where the model is most useful.

What is GPT-5.4 Mini?

GPT-5.4 Mini is a text-generating reasoning model with image understanding and support for external tools. In practical terms, it can read a prompt, work through a multi-step problem, inspect supported image inputs, produce code or written responses, and call tools when an application makes those tools available.

It sits below the larger GPT-5.4 model in the same family. The trade-off is straightforward: GPT-5.4 Mini is intended to deliver lower latency and lower per-token pricing, while the full GPT-5.4 remains the more appropriate choice when the highest available reasoning quality is more important than throughput or cost. The supplied research does not provide a direct benchmark table, so claims about its relative quality should be treated as positioning rather than as a quantified performance guarantee.

Key specifications

SpecificationGPT-5.4 Mini
ProviderOpenAI
Release dateMarch 17, 2026
Model familyGPT-5.4
Model typeLightweight reasoning model
Context window400,000 tokens
Maximum output128,000 tokens
InputText and images
OutputText
Knowledge cutoffAugust 31, 2025
StreamingSupported
Structured outputsSupported

A token is a unit of text used for processing and billing; it may be a whole word, part of a word, punctuation, or another small text segment. A 400,000-token context window gives the model substantial room for long instructions, source material, conversation history, codebases, or tool results. It does not mean that every long prompt will receive equally strong analysis, and the context limit is shared by the input and the model's generated output according to the API's usage rules.

The 128,000-token maximum output is an upper API limit, not a recommendation to request extremely long answers. Most applications should set a practical output limit suited to the task to reduce latency and spending.

Pricing and cost trade-offs

The documented standard API prices are:

  • Input: $0.75 per 1 million tokens.
  • Cached input: $0.075 per 1 million tokens.
  • Output: $4.50 per 1 million tokens.

Cached-input pricing applies when eligible repeated prompt content is served through OpenAI's prompt-caching mechanism. It can make a meaningful difference for applications that repeatedly send a stable system prompt, tool definitions, schemas, or reference material. Actual costs still depend on how much content is cached and how many new input and output tokens each request uses.

OpenAI's pricing documentation also notes a 10% uplift for regional processing endpoints. Batch API billing is a separate consideration, and the supplied research identifies batch support but does not provide a separate price amount here. For a continuously interactive application, the standard token rates and response latency are usually the most relevant comparison.

Compared with a larger model, GPT-5.4 Mini's main economic advantage is the combination of lower token prices and faster expected response times. That advantage matters in classification, code-assistance, agent loops, and high-volume automation, where thousands or millions of requests can make even a modest per-request difference significant. The counterargument is that a cheaper model can be a poor bargain if it requires repeated retries, extensive verification, or escalation to a larger model.

Inputs, outputs, and modalities

GPT-5.4 Mini accepts text and image inputs and returns text. It can therefore analyze screenshots, diagrams, photographs, interface states, and other supported visual material alongside written instructions. Image understanding is different from image generation: GPT-5.4 Mini does not directly return generated images, audio, video, or music.

The model also does not support native audio or video inputs according to the supplied specifications. Applications that need speech recognition, audio conversation, video understanding, or media generation as the primary model capability should use a more appropriate specialized or multimodal option, potentially combining services when necessary.

Text output can include ordinary prose, source code, structured data, or tool-call-related content. Structured outputs are useful when an application needs responses that follow a defined schema rather than free-form text. They are especially relevant for extracting fields from documents, routing requests, returning machine-readable decisions, or passing predictable data between software components.

Reasoning and coding capabilities

GPT-5.4 Mini supports controllable reasoning effort values of none, low, medium, high, and xhigh. Reasoning effort controls let an application trade response time and computational work against the depth of processing appropriate for the task. A simple transformation may not need extensive reasoning, while a difficult debugging, planning, or multi-step analysis task may benefit from a higher setting.

These settings should not be interpreted as a guarantee that a higher value always produces a correct answer. They can increase work and latency, and important outputs still require testing or human review. The model's documented knowledge cutoff is August 31, 2025, so information after that date should not be assumed to be present in its built-in knowledge. Web search or another connected data source can provide newer information when the application enables it.

Coding is one of GPT-5.4 Mini's strongest intended use cases. It can generate code, explain existing code, suggest changes, help diagnose errors, and participate in tool-driven software workflows. Its long context window is useful for repositories, large configuration files, logs, and documentation, while image input can help it interpret screenshots of user interfaces or error messages.

For production coding systems, the model should be paired with tests, static analysis, permission controls, and review. Generated code can contain defects, misunderstand requirements, or make unsafe assumptions even when the explanation appears convincing.

Tools and API support

GPT-5.4 Mini supports streaming and tool use through OpenAI's Responses API. Streaming allows an application to display or process partial text as it is generated instead of waiting for the complete response. This can make interactive interfaces feel faster, although it does not reduce the total amount of model work or token billing.

The supplied documentation identifies support for a broad set of Responses API tools, including:

  • Web search for retrieving current information.
  • File search for finding relevant content in connected files.
  • Image generation as a tool available within a workflow, not as GPT-5.4 Mini's native output modality.
  • Code Interpreter for executing supported analysis or coding tasks.
  • Hosted shell, apply patch, and skills-oriented tools for software and agent workflows.
  • Computer use for interacting with supported computer interfaces.
  • Model Context Protocol, or MCP, for connecting external tools and services.
  • Tool search for discovering or selecting available tools.

Tool support does not mean the model independently has unrestricted access to the internet, a computer, files, or private systems. The application must provide the relevant tool, configure permissions, and handle the returned results. Tool calls should be logged and constrained, especially when they can modify files, send messages, access accounts, or perform other consequential actions.

Main strengths and limitations

Where GPT-5.4 Mini is strongest

  • Cost-sensitive reasoning: It offers reasoning controls and a lower price point for applications that need more than simple text completion.
  • Fast interactive workflows: Its lightweight positioning is intended to reduce latency compared with larger models.
  • Long-context work: The 400,000-token context window supports large prompts, documents, code collections, and tool histories.
  • Coding and agent tasks: Coding, computer use, subagents, tool calling, and automation are central use cases.
  • Image understanding: It can combine text instructions with image inputs while keeping text as its output format.
  • Application integration: Streaming, structured outputs, web search, and other Responses API tools support production workflows.

Important limitations

  • It does not natively accept audio or video inputs.
  • It produces text rather than images, audio, video, or other direct media output.
  • Fine-tuning is not supported according to the supplied model record.
  • The knowledge cutoff is August 31, 2025; current information requires web search or another external source.
  • Lower cost and higher speed involve a quality trade-off when compared with a larger, more capable model such as GPT-5.4.
  • Tool access depends on the application and its permissions; listed tools are not automatically available in every request.
  • Outputs may be inaccurate or overconfident, so code, retrieved information, and high-impact decisions require verification.

Best use cases

GPT-5.4 Mini is a good fit for high-volume coding assistants, software agents, customer-support or operations workflows that need structured results, document and image analysis, and subagents that perform bounded tasks inside a larger system. It is also suitable for computer-use workflows where the model must interpret a screen, decide on the next action, and use an approved tool.

Examples include reviewing a pull request, extracting fields from many documents, classifying incoming requests, generating a first draft of a code change, searching a knowledge base, summarizing tool results, or coordinating several small tasks in an agent pipeline. The lower input price becomes particularly useful when prompts contain repeated instructions or large reference material that qualifies for caching.

When to choose GPT-5.4 Mini

Choose GPT-5.4 Mini when response speed, request volume, and operating cost matter alongside competent reasoning and coding. It is especially attractive when the task can be checked automatically, when a workflow can escalate difficult cases, or when the model is one component in a tool-using system rather than the sole decision-maker.

Choose the larger GPT-5.4 instead when the task justifies higher cost or latency in exchange for stronger reasoning quality. A small, fast model is not automatically the right choice for complex research, difficult planning, delicate code changes, or tasks where mistakes are expensive.

Choose a model or service designed for native audio, video, or media generation when those modalities are central to the product. GPT-5.4 Mini can inspect images and use image generation through a supported tool, but its own documented input-output profile is text and image input with text output.

Overall, GPT-5.4 Mini is best understood as a practical middle ground: more capable and tool-oriented than a basic low-cost text model, but less focused on maximum reasoning quality than a larger flagship model. Its value comes from applying that balance consistently across many responsive, verifiable workflows.


Answers to Frequently Asked Questions

When should I choose GPT-5.4 Mini instead of GPT-5.4?
Choose GPT-5.4 Mini when speed, request volume, and operating cost matter alongside capable reasoning and coding. It is well suited to high-volume automation, coding assistants, structured document or image analysis, subagents, and tool-driven workflows. Choose the larger GPT-5.4 when maximum reasoning quality is more important than latency or cost, especially for complex research, planning, or high-risk tasks.
What tools and modalities does GPT-5.4 Mini support?
GPT-5.4 Mini accepts text and image inputs and returns text. Through the Responses API, it supports streaming and tools such as web search, file search, Code Interpreter, computer use, hosted shell, apply patch, MCP, tool search, and image generation within a workflow. It does not natively accept audio or video or directly generate images, audio, or video.
How much does GPT-5.4 Mini cost through the API?
The standard API pricing is $0.75 per 1 million input tokens, $0.075 per 1 million cached input tokens, and $4.50 per 1 million output tokens. OpenAI also notes a 10% uplift for regional processing endpoints. Actual costs depend on token usage, caching eligibility, endpoint configuration, and billing method.
What is GPT-5.4 Mini?
GPT-5.4 Mini is OpenAI’s lightweight reasoning model for coding, computer-use agents, subagents, tool-enabled workflows, and high-volume applications. It supports text and image inputs, generates text, and is designed to offer lower latency and lower token costs than the larger GPT-5.4 model.
What are GPT-5.4 Mini’s context window and maximum output limits?
GPT-5.4 Mini has a 400,000-token context window and supports a maximum output of 128,000 tokens. The context window includes the request and generated output according to the API’s usage rules, while the maximum output is an upper limit rather than a recommended length for typical responses.


Sources 6
Provider

About OpenAI