What is GPT-5 nano?
GPT-5 nano is the smallest model in OpenAI's original GPT-5 API family. OpenAI designed it for applications that need many responses quickly and at low cost, rather than the deepest reasoning available from a larger model. Its main roles are classification, summarization, extraction, ranking, routing and other focused tasks where the input and expected output can be described clearly.
The model is available through OpenAI's Responses API and Chat Completions API. It can use reasoning tokens internally, and applications can configure reasoning effort to make a practical trade-off between response quality, latency and token usage. In simple terms, GPT-5 nano is intended to handle a large number of relatively narrow tasks without using a more expensive general-purpose model for every request.
Where GPT-5 nano fits in OpenAI's lineup
GPT-5 nano occupies the speed-and-cost end of the original GPT-5 family. Its positioning is different from a larger model intended for difficult reasoning, complex coding or nuanced research. The smaller model can be useful as the first stage of a routing system: it handles routine requests, while an application sends ambiguous or demanding cases to a more capable model.
This distinction matters because a low price does not mean that GPT-5 nano is the best choice for every prompt. A classification pipeline processing millions of short records may benefit from its economics, whereas a high-stakes analysis workflow may need a larger model and additional review. Readers evaluating the broader GPT-5 range can also compare it with the GPT-5 Mini, while OpenAI lists GPT-5.6 Luna as the recommended replacement for the deprecated dated snapshot.
Inputs, outputs and supported capabilities
GPT-5 nano accepts text and images and returns text. Its image input allows basic visual classification, screenshot interpretation, document understanding and image-assisted extraction. It does not natively generate images, audio or video, so applications that begin with speech, sound or video need a separate model or preprocessing step before sending suitable text or selected images to GPT-5 nano.
- Input: Text and images.
- Output: Text only.
- Context window: 400,000 tokens.
- Maximum output: 128,000 tokens.
- Reasoning: Configurable reasoning effort and reasoning tokens.
- Tools: Function calling and tool use.
- Response handling: Streaming and structured outputs.
- Efficiency features: Prompt caching and Batch API support.
A token is a small unit of text used for processing and billing. The 400,000-token context window is large enough for substantial documents or retrieved material, although sending more context still increases input usage and may not improve a poorly specified task. The 128,000-token maximum output is a documented ceiling, not a recommendation to request extremely long answers. For most classification, extraction and summarization jobs, shorter outputs are cheaper and faster.
Pricing and availability
The documented standard price is $0.05 per 1 million input tokens and $0.40 per 1 million output tokens. Cached input is priced at $0.005 per 1 million tokens. Cached pricing can be useful when an application repeatedly sends the same large instructions or reference material, while Batch API processing is relevant for asynchronous workloads that do not require an immediate response.
OpenAI released GPT-5 nano on August 7, 2025. The dated snapshot gpt-5-nano-2025-08-07 is marked deprecated but remains accessible according to the supplied research. OpenAI has announced API shutdown for that snapshot on December 11, 2026 and identifies GPT-5.6 Luna as the recommended replacement. Teams starting a new integration should therefore verify the currently supported model identifier and migration guidance rather than assuming that the dated snapshot will remain available indefinitely.
Reasoning, coding and tool use
GPT-5 nano supports reasoning, but its role is lightweight reasoning rather than maximum-depth problem solving. Configurable reasoning effort lets developers decide whether a task needs a faster response or additional internal work. This is useful for tasks such as assigning an intent label, selecting a route, extracting fields from a document or producing a short explanation from supplied material.
The model also supports function calling and tool use. A function call lets the model request an operation defined by the application, such as looking up a customer record, checking a rule or submitting structured data. The model does not itself replace the application's business logic: the software must validate the request, run the function and return the result. Structured outputs can make this exchange more reliable when downstream code expects a defined schema.
GPT-5 nano can assist with lightweight coding and coding subagents, particularly for code classification, small transformations, routing, documentation or tightly bounded implementation tasks. It is less appropriate for large architectural changes, difficult debugging across a complex repository or software engineering that depends on sustained, nuanced reasoning. Those cases should be evaluated with a larger coding or general-purpose model.
Best use cases
GPT-5 nano is most attractive when the workload is repetitive, high-volume and reasonably well specified. Typical examples include:
- Classification: Categorizing support messages, documents, products or moderation queues.
- Summarization: Creating short summaries of large collections of messages, reports or records.
- Extraction: Converting invoices, forms, emails or documents into structured fields.
- Routing: Detecting intent and choosing the next workflow, agent or knowledge source.
- Ranking: Filtering or ordering candidates before a more expensive review stage.
- Image-assisted processing: Identifying visual categories, interpreting screenshots or extracting simple information from images.
- Background processing: Running asynchronous jobs through caching or batch workflows.
- Lightweight coding support: Performing bounded code transformations or acting as a low-cost subagent.
For example, a help-desk system could use GPT-5 nano to classify incoming requests and extract account or issue fields. Only unusual or high-impact requests would then be escalated to a larger model or a human reviewer. This kind of two-stage design can reduce cost while reserving deeper reasoning for the cases that need it.
Limitations and trade-offs
The central trade-off is capability versus speed and cost. GPT-5 nano is cheaper and faster than a larger reasoning model, but it is not intended to deliver the same performance on ambiguous, multi-step or technically demanding problems. It may be a poor fit for complex software engineering, nuanced research, difficult mathematical reasoning, high-stakes decisions or tasks requiring consistently strong interpretation of complicated visual material.
Its documented knowledge cutoff is May 31, 2024. Web search, retrieval systems or user-provided documents can supply newer information during a request, but they do not change the model's underlying cutoff. Retrieved information should still be checked for relevance and accuracy, especially when the output affects legal, financial, medical or operational decisions.
GPT-5 nano also has no native audio or video input or output. An application handling a meeting recording, phone call or video may need transcription, audio analysis or frame extraction before GPT-5 nano can process the resulting text or images. Similarly, it cannot directly return generated images, audio or video.
When to choose GPT-5 nano
Choose GPT-5 nano when low unit cost, fast responses and throughput are more important than maximum reasoning depth. It is a sensible candidate when:
- the task can be expressed with clear instructions and a defined output;
- the application processes many requests;
- the output can be validated with a schema, rules or a later review step;
- text or image understanding is needed but native media generation is not;
- batch processing, prompt caching or streaming can improve the workflow; or
- a larger model would be excessive for routine decisions.
Use a larger alternative when the request is open-ended, difficult to verify, safety-critical or dependent on complex reasoning. A routing strategy can combine both approaches: GPT-5 nano handles ordinary cases, while confidence checks, validation rules or special conditions escalate difficult cases. Because the dated GPT-5 nano snapshot has a stated shutdown date, new systems should test the recommended successor before committing to a long-term deployment.
Implementation considerations
Applications should request structured outputs when the response feeds another service, validate all model-generated fields, and keep tool permissions narrowly scoped. Streaming can reduce perceived waiting time for interactive interfaces, while Batch API processing and prompt caching can improve economics for suitable asynchronous or repetitive workloads.
Before production release, test representative examples rather than relying only on the model's general description. Include difficult classifications, incomplete documents, unusual images and cases that should be escalated. The supplied editorial scores rate GPT-5 nano highly for speed and cost and more moderately for reasoning and coding; these are comparative editorial judgments, not OpenAI-published benchmark results. The verified specifications are its documented modalities, limits, pricing, supported interfaces and lifecycle information.

