What is GPT-5.4 nano?
GPT-5.4 nano is OpenAI's smallest and least expensive model in the GPT-5.4 family. It is an API model intended for applications that need many fast, relatively inexpensive model calls rather than the strongest available performance on difficult reasoning or coding problems.
OpenAI positions it for tasks such as classification, data extraction, ranking, routing, lightweight coding subagents, and other supporting steps in larger workflows. For example, an application might use GPT-5.4 nano to label incoming documents, extract fields from screenshots, filter search results, or decide which request should be sent to a more capable model.
The model is available through the OpenAI API and is not presented in the supplied research as a standalone ChatGPT model. Its canonical API identifier is gpt-5.4-nano; OpenAI also lists the dated snapshot gpt-5.4-nano-2026-03-17.
Where GPT-5.4 nano fits in the GPT-5.4 family
GPT-5.4 nano is the cost- and latency-oriented member of the family. The supplied research describes it as smaller and cheaper than the standard GPT-5.4 and GPT-5.4 mini, with lower expected performance on the hardest reasoning and coding workloads. The trade-off is deliberate: a system can use nano for routine subtasks while reserving a larger model for decisions that require more depth.
For a higher capability ceiling, readers can compare it with GPT-5.4 or GPT-5.4 mini. The supplied research does not provide a full price or benchmark table for those alternatives, so the useful distinction here is model role rather than an unsupported numerical comparison. GPT-5.4 nano is best understood as a fast worker or subagent, not as the default choice for every difficult task.
Context window and maximum output
GPT-5.4 nano has a 400,000-token context window. The context window is the amount of combined input and conversational or tool-provided material the model can consider in a request. This gives developers room to provide long documents, multiple files, image-related context, or substantial intermediate results without immediately splitting the task into many smaller requests.
The maximum output length is 128,000 tokens. That limit is much larger than normally required for classification or extraction, but it can be useful for workflows that generate long structured results or perform extended reasoning. A large maximum does not mean every request should ask for a long response: shorter output limits generally make applications easier to control and can reduce unnecessary usage.
Pricing and availability
OpenAI's documented standard pricing is:
| Token type | Price per 1 million tokens |
|---|---|
| Input tokens | $0.20 |
| Cached input tokens | $0.02 |
| Output tokens | $1.25 |
Cached input pricing applies when eligible prompt content can be reused through prompt caching. Output tokens cost more than ordinary input tokens, so applications that generate concise labels, extracted fields, or routing decisions can benefit from nano's economics particularly clearly.
The model supports the Batch API, which has separate batch pricing rules. The supplied research also notes that regional processing endpoints carry a 10% pricing uplift. Developers should therefore distinguish the standard token prices from any batch, regional-processing, or other deployment-specific price shown for their account.
GPT-5.4 nano was released on March 17, 2026 and is listed as currently available in the supplied research.
Supported inputs and outputs
GPT-5.4 nano accepts text and image inputs and returns text. Image input allows the model to interpret visual material such as screenshots, scanned pages, charts, or photographs when those images are supplied through a supported API workflow.
- Text input: Supported.
- Image input: Supported for visual understanding.
- Text output: Supported.
- Audio input: Not supported.
- Video input: Not supported.
- Native image, audio, video, speech, music, or embedding output: Not supported.
This distinction matters when evaluating tool-enabled applications. The Responses API can expose tools such as image generation, but a tool invoked during a request is not the same as GPT-5.4 nano natively producing an image. The model itself remains a text-output model with text and image understanding on the input side.
Reasoning and coding capabilities
GPT-5.4 nano supports configurable reasoning effort levels of none, low, medium, high, and xhigh. The default is none. In practice, the setting lets developers choose between lower latency for straightforward work and more deliberate processing for tasks that benefit from additional reasoning.
A none setting is suitable for predictable operations such as assigning a category, extracting a known set of fields, or ranking items against a simple rubric. Higher settings may be more appropriate when a request involves several constraints, ambiguous evidence, or a lightweight coding task. More reasoning can affect response time and usage, so it should be applied selectively rather than automatically to every request.
The supplied research characterizes GPT-5.4 nano's coding capability as useful for lightweight coding subagents. It is therefore a reasonable candidate for code classification, small transformations, test triage, simple snippets, or delegated repository tasks with narrow scope. It is not the preferred option for the most difficult autonomous coding, long-horizon implementation, or complex debugging work. Those use cases are better candidates for a larger GPT-5.4-family model, such as GPT-5.4 Pro, when its higher capability ceiling justifies the additional cost or latency.
Tools and API features
GPT-5.4 nano supports function calling and tool use, allowing an application to give the model access to defined actions or external services. It also supports structured outputs, which help developers request responses that follow a specified machine-readable structure instead of relying only on free-form text.
- Function calling and tool use
- Web search through supported OpenAI API tooling
- File search
- Code interpreter
- Hosted shell
- Skills and MCP integrations
- Image-generation tools through supported workflows
- Streaming responses
- Prompt caching
- Batch API processing
Streaming sends partial output as it becomes available, which can improve the perceived responsiveness of an application. Batch processing is useful when large numbers of requests can be handled asynchronously. Structured outputs are especially useful for extraction and classification because downstream software can validate fields instead of parsing an informal paragraph.
Web search can provide newer information during a request, but it does not change the model's underlying knowledge cutoff. The official cutoff specified in the supplied research is August 31, 2025.
Main strengths and trade-offs
GPT-5.4 nano's strongest advantage is the combination of low token pricing, fast expected response behavior, a large context window, and broad API support. It can handle both text and image understanding while producing structured text, calling tools, and operating in high-volume workflows. That combination makes it more useful than a text-only, single-purpose classifier when an application must process varied documents or connect model decisions to software actions.
Its main trade-off is capability. OpenAI's positioning places nano below larger GPT-5.4-family models for the most demanding reasoning and coding tasks. A cheaper model can also create hidden costs if it makes enough mistakes that requests must be retried or reviewed manually. The right comparison is therefore not only price per token, but total workflow cost, including accuracy requirements, latency, verification, and escalation to a larger model.
Best use cases for GPT-5.4 nano
- High-volume classification: Categorizing support tickets, documents, transactions, or user requests.
- Data extraction: Turning text, forms, or images into structured fields.
- Ranking and filtering: Ordering candidates, search results, or records against a defined rubric.
- Image understanding: Quickly interpreting screenshots, scanned material, or other supplied images.
- Routing: Deciding which workflow, tool, queue, or larger model should handle a request.
- Lightweight coding subagents: Performing narrow code transformations, triage, or other bounded programming tasks.
- Delegated agent subtasks: Handling repetitive steps inside a larger multi-model system.
- Cost-sensitive tool calling: Making structured decisions before an application invokes external services.
When to choose GPT-5.4 nano
Choose GPT-5.4 nano when request volume is high, the task can be clearly specified, and low latency or low cost matters more than maximum answer quality. It is particularly attractive when the output can be validated automatically, such as a fixed list of labels, a schema of extracted fields, a ranking score, or a tool-call decision.
It is also a practical choice for a two-stage architecture: nano handles routine screening or preparation, and a larger model reviews only ambiguous or high-value cases. This approach can control costs without forcing the most capable model to process every request.
Choose a larger model instead when the task requires sustained autonomous reasoning, difficult coding, nuanced judgment, or the highest available reliability. GPT-5.4 nano is also unsuitable when the application needs native audio, video, or image generation, fine-tuning, or the computer-use tool. The supplied research specifically identifies fine-tuning and computer use as unsupported.
Limitations to consider
GPT-5.4 nano does not support fine-tuning or computer use. It cannot directly return audio, video, or image output, and it does not accept audio or video inputs. Tool availability should not be mistaken for native support of every modality exposed by the surrounding API.
Its knowledge cutoff is August 31, 2025. Web search, file search, user-provided documents, and other tools can add current information to a request, but developers should still verify important results and account for tool failures or incomplete sources. The model's low price also does not guarantee correctness; applications handling consequential decisions should use validation, human review, or escalation rules.
Bottom line
GPT-5.4 nano is a low-cost, high-throughput API model for narrow tasks that benefit from speed, structured responses, image understanding, and tool access. Its 400,000-token context window, 128,000-token maximum output, configurable reasoning, and support for caching and batch processing give it room to serve as more than a basic classifier. Its best role is a fast worker or subagent. For the hardest reasoning, complex coding, computer-use workflows, fine-tuning, or native non-text generation, another option is more appropriate.

