GPT-4.1

GPT-4.1 nano

by OpenAI · Deprecated; currently accessible as of September 23, 2026; scheduled for shutdown on October 23, 2026

GPT-4.1 nano is OpenAI’s lightweight GPT-4.1 model for fast, inexpensive, high-volume API workloads. It supports text and image input, text output, a 1,047,576-token context window, function calling, structured outputs, streaming, caching, batch processing, and fine-tuning. Its main limitations are weaker reasoning and coding performance than larger models and a scheduled shutdown on October 23, 2026.

Text Reasoning Coding
GPT-4.1 nano is the smallest and least expensive member of OpenAI’s GPT-4.1 family. It is designed for applications that need to process many requests quickly and cheaply rather than perform deep reasoning on difficult problems. The model supports text and image input, produces text, and offers a million-token context window alongside function calling and structured outputs. Its low token prices make it attractive for classification, extraction, routing, summarization, document triage, and lightweight assistants. However, OpenAI has deprecated the model and scheduled its shutdown for October 23, 2026, so new deployments should account for migration work.
Outputs

What GPT-4.1 nano can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Fine-tuning Structured output Prompt caching Batch API
Model profile

Performance characteristics

4/10 Reasoning
5/10 Coding
10/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family GPT-4.1
Model type Lightweight
Context window 1.05M tokens
Maximum output 33K tokens
Knowledge cutoff 2024-06-01
Release date 2025-04-14
Status Deprecated; currently accessible as of September 23, 2026; scheduled for shutdown on October 23, 2026
Deprecation date 2026-04-22
Shutdown date 2026-10-23
Knowledge cutoff notes

OpenAI's exact model page lists June 1, 2024 as the knowledge cutoff. Web search or external retrieval can provide newer information during use but does not change the underlying cutoff.

Model notes

GPT-4.1 nano is the fastest and most cost-efficient GPT-4.1 variant and operates without a separate reasoning step. The canonical alias is gpt-4.1-nano and the dated snapshot is gpt-4.1-nano-2025-04-14. Standard inference supports a 1,047,576-token context window and 32,768-token maximum output. Fine-tuning documentation lists the dated snapshot with a 128,000-token inference context and 65,536-token example limit, so fine-tuned deployments may differ from the base endpoint. OpenAI documents text and image input, text output, function calling, structured outputs, streaming, prompt caching, Batch API access, and fine-tuning. The model does not natively produce images, audio, video, speech, music, or embeddings. OpenAI's current deprecation schedule lists October 23, 2026 as the shutdown date and GPT-5.6 Luna as the recommended replacement.

Cost

Model pricing

Input $0.10 per 1M input tokens; $0.025 per 1M cached input tokens
Output $0.40 per 1M output tokens
Model guide

GPT-4.1 nano: OpenAI’s Low-Cost Model for Fast, High-Volume Workloads

GPT-4.1 nano is OpenAI’s smallest GPT-4.1 model, built for fast, inexpensive, high-volume API workloads. It accepts text and images, generates text, supports a 1,047,576-token context window, function calling, structured outputs, streaming, prompt caching, batch processing, and fine-tuning. It is a practical choice for classification, extraction, routing, summarization, and lightweight assistants, but is less suitable for complex reasoning, demanding coding, or applications needing a long-term model because it is deprecated and scheduled to shut down on October 23, 2026.

What is GPT-4.1 nano?

GPT-4.1 nano is OpenAI’s smallest model in the GPT-4.1 family. OpenAI positions it for low-latency, high-volume API workloads where the application needs a useful text model at a very low token cost. It is a non-reasoning model: unlike models designed around an explicit reasoning step, GPT-4.1 nano does not perform a separate reasoning phase before producing its answer. That design helps keep responses fast, but it also makes the model a less suitable choice for difficult multi-step problems.

The model is available through OpenAI’s supported API interfaces, including the Responses API and Chat Completions API. Its primary role is not to replace a more capable model on every task. Instead, it can handle routine requests in a larger system, leaving difficult or ambiguous cases to a stronger model.

According to the supplied model information, GPT-4.1 nano remained accessible as of September 23, 2026, but was deprecated and scheduled to shut down on October 23, 2026. That lifecycle status is important: the model may still be useful for maintaining an existing application in the short term, but it is a poor foundation for a new long-lived deployment unless migration has already been planned.

Specifications and limits

SpecificationGPT-4.1 nano
ProviderOpenAI
Model IDgpt-4.1-nano
Model familyGPT-4.1
Release dateApril 14, 2025
Knowledge cutoffJune 1, 2024
Context window1,047,576 tokens
Maximum output32,768 tokens
InputText and images
OutputText
Reasoning designNon-reasoning model without a separate reasoning step

The context window is the amount of text and other supported input the model can consider in one request. At 1,047,576 tokens, GPT-4.1 nano can process unusually large prompts, long documents, or substantial collections of context at once. This does not guarantee that every detail will receive equal attention, and large requests can still increase processing time and cost.

The 32,768-token maximum output is a ceiling rather than a requirement. Most classification, extraction, routing, and summarization tasks should use much shorter responses. The model’s knowledge cutoff is June 1, 2024; applications requiring newer information must supply current data through their own retrieval or tools.

Pricing and cost profile

Standard pricing is listed per 1 million tokens:

  • Input: $0.10 per 1 million tokens
  • Cached input: $0.025 per 1 million tokens
  • Output: $0.40 per 1 million tokens

Cached input pricing applies when eligible prompt content can be reused, which is useful for applications that repeatedly send the same instructions, policies, schemas, or reference material. OpenAI also supports GPT-4.1 nano through the Batch API, which provides a 50% discount on standard token prices according to the supplied research. Batch processing is more appropriate for workloads that do not require an immediate response.

The pricing structure favors applications that make many small or moderate requests. For example, a service that classifies incoming messages, extracts fields from documents, assigns support tickets to queues, or produces short summaries can keep per-request costs low. Output is priced higher than input, so applications can reduce spending by requesting concise, structured responses rather than long explanations.

Input, output, and tool support

Text and image input

GPT-4.1 nano accepts text and image input and produces text output. Image understanding can support lightweight visual classification, document extraction, chart interpretation, and image-assisted routing. The documented interface does not support audio or video input, and the model does not natively generate images, audio, video, speech, music, or embeddings.

Function calling and structured outputs

The model supports function calling, allowing an application to describe external operations that the model can request. A customer-service workflow, for example, could let GPT-4.1 nano identify an operation such as looking up an order or creating a ticket, while the application performs the actual action.

GPT-4.1 nano also supports structured outputs. This allows an application to request data that follows a defined schema, such as a JSON object containing a category, confidence field, extracted names, and routing decision. Structured outputs are especially useful for extraction and automation because downstream software does not have to interpret a free-form paragraph. The supplied research does not separately verify a distinct legacy JSON-mode capability, so structured outputs should not automatically be treated as the same feature.

Streaming, caching, batch processing, and fine-tuning

OpenAI documents streaming, prompt caching, Batch API access, and fine-tuning support for GPT-4.1 nano. Streaming can make a response appear sooner by delivering generated text incrementally, while caching and batch processing can reduce costs in suitable workloads.

Fine-tuning documentation identifies the dated snapshot gpt-4.1-nano-2025-04-14 as supporting supervised fine-tuning and related optimization methods. The supplied information warns that fine-tuning limits can differ from the standard base endpoint: fine-tuning documentation lists a 128,000-token inference context and a 65,536-token example limit for the fine-tuning configuration. Developers should therefore avoid assuming that the base model’s 1,047,576-token context window applies unchanged to a fine-tuned deployment.

Main strengths and trade-offs

  • Low cost: The standard input and output prices are among the model’s clearest advantages for high-volume processing.
  • Fast responses: Its lightweight, non-reasoning design is intended to reduce latency.
  • Very large context: The standard endpoint supports a 1,047,576-token context window.
  • Useful automation features: Function calling, structured outputs, streaming, caching, batch processing, and fine-tuning support make it practical in application workflows.
  • Image understanding: Text and image input allow it to handle more than text-only classification and extraction.

The trade-off is capability. GPT-4.1 nano is not designed for deep reasoning, difficult mathematics, complex research, or demanding software engineering. It may produce a quick and inexpensive answer where a larger model would spend more resources checking assumptions, handling edge cases, or solving a complicated chain of dependencies.

The supplied evaluation data rates its speed and cost most strongly, while its reasoning and coding scores are lower. These are editorial or database evaluations, not provider-published benchmark results, and should be treated as directional rather than as formal guarantees.

Best use cases for GPT-4.1 nano

GPT-4.1 nano is most appropriate when the task is repeatable, well-defined, and sensitive to latency or operating cost. Suitable examples include:

  • Classifying messages, documents, support requests, or product records
  • Extracting fields from text or images into a defined schema
  • Tagging, moderation support, and content routing
  • Summarizing large volumes of routine material
  • Document triage and image-assisted classification
  • Lightweight customer-service automation
  • Simple tool-calling workflows
  • High-volume preprocessing before a more capable model handles selected cases

A practical architecture can use GPT-4.1 nano as a first-pass filter. It might identify straightforward requests, extract structured facts, or route uncertain cases. More capable models can then handle the smaller set of requests that require detailed reasoning or nuanced judgment.

When to choose this model

Choose GPT-4.1 nano when low cost and quick responses matter more than maximum reasoning or coding ability. It is particularly compelling when each request is short, the desired output has a predictable structure, and the application may issue thousands or millions of calls. Image input, function calling, and structured outputs also make it more useful than a basic text-only classifier for workflow automation.

Consider another option when the task involves advanced software engineering, difficult research, complex mathematical reasoning, high-stakes decision support, or subtle instructions that require extensive deliberation. A larger or newer model may cost more and respond more slowly, but that trade-off can be justified when errors are expensive.

GPT-4.1 nano’s deprecation is an additional reason to evaluate alternatives. The supplied deprecation information identifies GPT-5.6 Luna as the recommended replacement. Any migration should be tested rather than assumed to be drop-in compatible. Compare output style, tool-call behavior, structured-output conformance, image understanding, latency, token use, and total cost before switching production traffic.

Availability and migration considerations

The canonical model ID is gpt-4.1-nano, and OpenAI also lists the dated snapshot gpt-4.1-nano-2025-04-14. The snapshot and canonical model are subject to lifecycle differences, so applications should record which identifier they use and monitor OpenAI’s current model documentation.

Because shutdown is scheduled for October 23, 2026, new projects should not treat GPT-4.1 nano as a permanent dependency without a tested replacement. Existing users should inventory prompts, schemas, tool definitions, image inputs, fine-tuned versions, and batch jobs before migration. The model remains attractive as a description of the speed-and-cost trade-off OpenAI offered in the GPT-4.1 family, but its scheduled retirement makes continuity a more important concern than it would be for an actively supported model.


Answers to Frequently Asked Questions

Is GPT-4.1 nano still available, and should new projects use it?
According to the supplied information, GPT-4.1 nano was available as of September 23, 2026, but was scheduled for shutdown on October 23, 2026. Existing applications may use it temporarily, but new long-lived projects should evaluate a replacement and test compatibility for prompts, tools, structured outputs, image inputs, latency, and cost.
Does GPT-4.1 nano support images, function calling, and structured outputs?
Yes. GPT-4.1 nano accepts text and image input, produces text output, supports function calling, and can generate structured outputs that follow a defined schema. It does not natively support audio or video input or generate images, audio, video, speech, music, or embeddings.
What are the context window and output limits of GPT-4.1 nano?
GPT-4.1 nano has a standard context window of 1,047,576 tokens and a maximum output of 32,768 tokens. Its knowledge cutoff is June 1, 2024, so applications needing current information must provide updated data through retrieval or external tools.
What is GPT-4.1 nano best used for?
GPT-4.1 nano is best suited to fast, repeatable, high-volume tasks such as text and image classification, document and data extraction, moderation support, ticket routing, short summaries, lightweight customer service, and simple tool-calling workflows.
How much does GPT-4.1 nano cost?
Standard pricing is $0.10 per 1 million input tokens, $0.025 per 1 million cached input tokens, and $0.40 per 1 million output tokens. OpenAI’s Batch API can provide a 50% discount on standard token prices for workloads that do not require immediate responses.


Sources 6
Provider

About OpenAI