GPT-5

GPT-5.1

by OpenAI · Current

OpenAI's GPT-5.1 is an API model for coding, long-context analysis, and agentic workflows. It supports configurable reasoning, a 400,000-token context window, 128,000-token maximum output, text and image input, function calling, structured outputs, streaming, batch processing, and prompt caching.

Text Reasoning Coding
GPT-5.1 is OpenAI's general-purpose model for coding, multi-step reasoning, and applications that use tools or structured workflows. Released in the API on November 13, 2025, it accepts text and images, returns text, supports reasoning settings from none to high, and provides a 400,000-token context window with up to 128,000 output tokens.
Outputs

What GPT-5.1 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family GPT-5
Model type General Purpose
Context window 400K tokens
Maximum output 128K tokens
Knowledge cutoff September 30, 2024
Release date November 13, 2025
Status Current
Knowledge cutoff notes

OpenAI's current model documentation explicitly lists September 30, 2024 as the knowledge cutoff. This is separate from information supplied through web search, retrieval, files, tools, or application context.

Model notes

GPT-5.1 supports reasoning effort values none, low, medium, and high, with none as the default. The exact model accepts text and image input and returns text output. It supports function calling, structured outputs, streaming, Batch API processing, and prompt caching, including extended retention of up to 24 hours when configured. The documented knowledge cutoff is September 30, 2024. Standard pricing is $1.25 per 1M input tokens, $0.125 per 1M cached input tokens, and $10 per 1M output tokens. Editorial scores are comparative estimates, not vendor-published ratings. The canonical dated snapshot is gpt-5.1-2025-11-13.

Cost

Model pricing

Input $1.25 per 1 million input tokens; $0.125 per 1 million cached input tokens
Output $10.00 per 1 million output tokens
Model guide

GPT-5.1: Features, Pricing, Context Window and API Details

GPT-5.1 is OpenAI's API model for coding, long-context analysis, structured responses, and agentic workflows. It combines configurable reasoning effort with a 400,000-token context window, 128,000-token maximum output, text and image input, function calling, structured outputs, streaming, batch processing, and prompt caching.

What is GPT-5.1?

GPT-5.1 is a general-purpose OpenAI API model designed for coding, agentic applications, long-context analysis, and tasks that require a controlled balance between reasoning depth, response speed, and cost. An agentic application is software that can plan steps, call functions or external tools, inspect results, and continue working toward a goal. GPT-5.1 is intended to serve as the language and reasoning component inside those workflows.

OpenAI released GPT-5.1 in the API on November 13, 2025. The standard model identifier is gpt-5.1, and the dated snapshot is gpt-5.1-2025-11-13. The model is separate from GPT-5.1 ChatGPT aliases and GPT-5.1 Codex variants, which have their own identities and lifecycle information.

Its strongest fit is not a simple chat application that only needs short answers. GPT-5.1 is more useful when an application needs reliable structured responses, code generation or revision, image-aware analysis, tool calls, and enough context to work across large documents or multi-step tasks.

GPT-5.1 specifications at a glance

SpecificationGPT-5.1
ProviderOpenAI
API releaseNovember 13, 2025
Context window400,000 tokens
Maximum output128,000 tokens
InputText and images
OutputText
Reasoning effortNone, low, medium, or high
Knowledge cutoffSeptember 30, 2024
Fine-tuningNot supported

A token is a unit used to measure text for model processing and billing; it may represent a whole word, part of a word, punctuation, or other text. The 400,000-token context window is the total amount of input and generated context the model can handle in a request, subject to the applicable API behavior and output limit. The 128,000-token maximum output is a separate ceiling for the response.

Reasoning and coding capabilities

GPT-5.1 provides four reasoning-effort settings: none, low, medium, and high. The default is none. Lower settings are intended for tasks where quick responses matter, such as routine transformations, straightforward extraction, or ordinary code assistance. Higher settings allocate more reasoning to difficult problems, which can help with complex planning, debugging, analysis, and multi-step agentic work.

The trade-off is that additional reasoning can increase latency and token consumption. A high setting is therefore not automatically the best choice for every request. Applications that process many simple requests may prefer no reasoning or low reasoning, while difficult repository-level changes or complicated plans may justify medium or high reasoning.

OpenAI positions GPT-5.1 particularly strongly for coding and agentic tasks. It can generate new code, revise existing code, interpret complex technical instructions, and participate in workflows where an application retrieves information or invokes functions between model responses. The supplied research does not provide benchmark scores, so performance claims should be understood as product positioning and practical capability descriptions rather than independently verified rankings.

Input, output, and modality support

GPT-5.1 accepts text and image input and produces text output. Image input can be useful for analyzing screenshots, diagrams, visual documents, or other image-based information alongside written instructions. The exact model does not natively accept audio or video input, and it does not generate images, audio, or video.

This distinction matters when selecting a model for a larger application. GPT-5.1 can be part of a multimodal workflow because it understands images, but it is not a direct media-generation model. An application requiring native speech processing, video understanding, image generation, or audio generation would need another model or an additional processing component.

The 400,000-token context window makes GPT-5.1 suitable for large prompts, extended code context, long documents, and multi-step conversations that would exceed the capacity of smaller-context options. A large context does not guarantee that every detail will be equally important or accurately interpreted, so retrieved material should still be selected and organized carefully.

API and developer features

GPT-5.1 is available through both the Responses API and the Chat Completions API. It supports streaming, which lets an application receive a response progressively instead of waiting for the complete answer. This can improve the perceived responsiveness of interactive applications, although it does not eliminate the model's underlying processing time.

Function calling allows the model to request that an application execute a defined function, such as querying a database, creating a ticket, or running an internal operation. The application remains responsible for executing and validating that action. GPT-5.1 also supports structured outputs, allowing developers to constrain responses to a specified schema for uses such as data extraction, classification, and workflow handoffs.

The model supports the Batch API for eligible asynchronous workloads. Batch processing is useful when results do not need to be returned immediately and can provide a different cost and throughput option from interactive requests. Prompt caching is also supported, including extended retention of up to 24 hours when the applicable API setting is configured. Caching can reduce repeated-input costs and processing overhead for applications that reuse long prompt prefixes.

GPT-5.1 can be used with OpenAI's first-party web-search tool through supported Responses API workflows. Web search supplies externally retrieved information during a request; it does not change the model's underlying knowledge cutoff.

GPT-5.1 pricing

OpenAI's standard pricing for GPT-5.1 is:

  • Input: $1.25 per 1 million tokens
  • Cached input: $0.125 per 1 million tokens
  • Output: $10.00 per 1 million tokens

Input and output are billed separately. Output tokens are substantially more expensive than ordinary input tokens, so applications can control costs by avoiding unnecessarily long responses, selecting an appropriate reasoning setting, and using structured prompts that produce only the information needed. Repeated prompt content may qualify for the cached-input rate when prompt caching is configured and applicable.

The Batch API has separate discounted pricing and asynchronous processing for eligible workloads. The supplied pricing information does not provide a single all-purpose batch rate, so batch costs should be checked against OpenAI's current pricing documentation before implementation.

Knowledge cutoff and limitations

The documented knowledge cutoff for GPT-5.1 is September 30, 2024. This means the model's built-in knowledge does not by itself cover later events or information. Applications that need current facts should supply retrieved context, use supported web search, or connect the model to an appropriate data source.

GPT-5.1 is not fine-tunable according to the supplied model information. It also lacks native audio and video input and cannot directly produce images, audio, or video. These limitations make it a less suitable choice for applications centered on speech-to-speech interaction, video analysis, or media generation.

Higher reasoning effort can increase both response time and token usage. A larger context window can also increase input costs when applications send extensive material on every request. Finally, structured outputs help enforce a response format but do not make the underlying information automatically correct; applications should still validate extracted values and tool arguments.

When to choose GPT-5.1

Choose GPT-5.1 when the application needs a combination of coding ability, long context, configurable reasoning, and tool-oriented workflow support. It is a strong candidate for:

  • Generating, debugging, refactoring, and reviewing substantial codebases
  • Agentic workflows that combine planning with function or tool calls
  • Long-document or repository analysis with text and image inputs
  • Structured extraction into a JSON schema or another defined application format
  • Tasks where the application needs to tune the balance between latency, reasoning depth, and cost
  • Applications that benefit from streaming, prompt caching, or asynchronous batch processing

GPT-5.1 is less appropriate when the main requirement is native audio or video processing, direct image or media generation, or fine-tuning. A smaller or faster model may be more economical for high-volume, simple classification or transformation tasks. Conversely, a model with stronger specialized performance may be preferable if a workload has a narrow requirement that GPT-5.1 does not address, although the supplied research does not establish benchmark-based superiority for any particular alternative.

For current information, GPT-5.1 should be paired with retrieval or web search rather than treated as a live source of facts. For safety-critical, financial, legal, or operational decisions, its output should be checked by suitable validation and human review.

Bottom line

GPT-5.1 is an API-focused OpenAI model for demanding coding, analysis, and agentic workflows. Its defining combination is a 400,000-token context window, up to 128,000 output tokens, configurable reasoning from none to high, image-aware input, text output, structured responses, function calling, streaming, batch processing, and prompt caching.

Its main decision trade-off is capability versus latency and cost. Use lower reasoning for speed-sensitive routine work and higher reasoning for complex planning or debugging. Select GPT-5.1 when those controls and developer features matter; choose another option when the workload requires native audio or video, media generation, fine-tuning, or a simpler low-cost model for straightforward tasks.


Answers to Frequently Asked Questions

How much does GPT-5.1 cost?
GPT-5.1 costs $1.25 per 1 million input tokens, $0.125 per 1 million cached input tokens, and $10.00 per 1 million output tokens. Input and output are billed separately, and eligible Batch API workloads use separate discounted pricing.
What is GPT-5.1 and what is it used for?
GPT-5.1 is an OpenAI API model designed for coding, long-context analysis, structured responses, and agentic applications that use tools or function calls. It is especially suitable for multi-step workflows, repository-level coding tasks, document analysis, and applications that need configurable reasoning.
What reasoning settings does GPT-5.1 support?
GPT-5.1 supports four reasoning-effort settings: none, low, medium, and high. The default is none. Lower settings generally provide faster and less expensive responses, while higher settings can improve performance on complex planning, debugging, analysis, and multi-step agentic tasks.
What is GPT-5.1's context window and maximum output size?
GPT-5.1 has a 400,000-token context window and supports a maximum output of 128,000 tokens. The context window covers the information processed in a request, while the output limit applies only to the generated response.
What are GPT-5.1's main limitations?
GPT-5.1 has a knowledge cutoff of September 30, 2024, is not fine-tunable, accepts text and image inputs but not native audio or video, and produces text output only. It cannot directly generate images, audio, or video, and current information should be supplied through retrieval, web search, or another connected data source.


Sources 7
Provider

About OpenAI