GPT-5.1-Codex

GPT-5.1-Codex-Max

by OpenAI · Retired; API access ended 2026-07-23

GPT-5.1-Codex-Max was OpenAI’s specialized model for long-running software-engineering tasks. It supported a 400,000-token context window, 128,000-token maximum output, image input, reasoning, function calling, streaming, structured outputs, and context compaction across multiple windows. It was retired from API access on July 23, 2026.

Text Reasoning Coding
GPT-5.1-Codex-Max was built for software-engineering tasks that take many steps, involve multiple files, and require an agent to retain project context while iterating. OpenAI positioned it for repository-scale refactoring, debugging, code review, pull-request creation, frontend development, and other extended Codex workflows. Its unusually large context window and compaction support made it suitable for long-running tasks, but its API retirement means it is no longer a current choice for new production integrations.
Outputs

What GPT-5.1-Codex-Max can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Structured output Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
10/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-5.1-Codex
Model type Coding
Context window 400K tokens
Maximum output 128K tokens
Knowledge cutoff 2024-09-30
Release date 2025-11-19
Status Retired; API access ended 2026-07-23
Deprecation date 2026-04-22
Shutdown date 2026-07-23
Knowledge cutoff notes

The official model documentation lists September 30, 2024 as the knowledge cutoff. Web search, tools, or external context do not change the underlying cutoff.

Model notes

GPT-5.1-Codex-Max was purpose-built for agentic coding and was available through the Responses API. OpenAI described it as the first model natively trained to operate across multiple context windows using compaction, enabling coherent tasks spanning millions of tokens. It supported text and image input, text output, reasoning tokens, function calling, streaming, and structured outputs. OpenAI announced deprecation on April 22, 2026 and ended API access on July 23, 2026. The deprecation documentation listed GPT-5.6 Sol as the recommended replacement. Editorial scores are comparative estimates, not official OpenAI ratings.

Cost

Model pricing

Input $1.25 per 1M tokens; cached input $0.125 per 1M tokens
Output $10.00 per 1M tokens
Model guide

GPT-5.1-Codex-Max: OpenAI’s Long-Running Agentic Coding Model

GPT-5.1-Codex-Max was OpenAI’s specialized model for long-running, agentic software-engineering work. It combined reasoning, coding, tool use, image input, a 400,000-token context window, 128,000-token maximum output, and context compaction across multiple windows. It was launched in Codex on November 19, 2025, later offered through the Responses API, and retired from API access on July 23, 2026.

What GPT-5.1-Codex-Max was designed to do

GPT-5.1-Codex-Max was a specialized OpenAI model for agentic coding. In this context, agentic coding means the model can work through a software task in multiple stages: inspect a repository, reason about the required changes, edit several files, use tools, review the result, and continue based on feedback.

Rather than focusing primarily on short conversational answers or isolated code completion, GPT-5.1-Codex-Max was optimized for sustained software-engineering work inside Codex. Its target tasks included multi-file implementation, repository-scale refactoring, deep debugging, code review, frontend development, pull-request creation, and technical question answering.

The model launched in Codex on November 19, 2025. It became available through OpenAI’s API in December 2025, using the Responses API. OpenAI later announced its deprecation on April 22, 2026, and API access ended on July 23, 2026. The deprecation documentation identified GPT-5.6 Sol as the recommended replacement, so GPT-5.1-Codex-Max should now be treated as a retired model rather than a new production option.

Long-running context and compaction

The model’s defining capability was support for work across multiple context windows. A context window is the amount of conversation, code, tool output, and other information the model can consider in one request. GPT-5.1-Codex-Max exposed a 400,000-token context window and supported a process OpenAI called compaction.

Compaction preserves the most relevant information from an earlier part of an agent session so work can continue in a later context window. This is important for large repositories and debugging sessions where the complete task history may eventually exceed the limit of one request. OpenAI described the model as capable of coherent work spanning millions of tokens across a single task through this multi-window workflow.

Compaction did not mean that every original token remained available in full detail forever. As with any context-management process, the agent still depended on retaining the information most relevant to the next stage. The practical advantage was that a long task could continue without restarting from scratch whenever one individual context window became full.

Coding and reasoning capabilities

GPT-5.1-Codex-Max was trained on software-engineering tasks as well as mathematics, research, computer use, and related reasoning domains. Its coding orientation made it particularly relevant when the desired result was a working change across a codebase rather than a short explanatory response.

  • Repository-scale refactoring and migration work
  • Multi-file implementation and debugging
  • Code review and pull-request creation
  • Frontend and application development
  • Extended agent loops using tools and iterative feedback
  • Technical question answering connected to software projects

Reasoning tokens were supported, allowing the model to spend additional computation on problems that required planning or analysis. The supplied research does not provide a single benchmark score or a guaranteed completion rate, so claims about reasoning quality should be understood as positioning and capability descriptions rather than a promise of a specific outcome.

API specifications and supported inputs

When it was available, GPT-5.1-Codex-Max was accessed through the Responses API. It accepted text and images as input and returned text. Image input could therefore provide visual context alongside code or instructions, but the model was not a native image, audio, or video generation system.

SpecificationValue
Context window400,000 tokens
Maximum output128,000 tokens
Knowledge cutoffSeptember 30, 2024
Input modalitiesText and image
Output modalityText
APIResponses API
ReasoningSupported
Function callingSupported
StreamingSupported
Structured outputsSupported

The 128,000-token maximum output applied to an individual model context. It should not be interpreted as a guarantee that every request could produce that much useful code or text. Output length also depended on the task, instructions, available context, and the agent’s tool workflow.

Tools, structured output, and agent workflows

Function calling allowed an application to expose tools that the model could request during a task. In a coding workflow, those tools could support actions such as inspecting files, running checks, or interacting with an external development system, subject to the application’s implementation and permissions. The model’s support for function calling was therefore useful for building an agent loop, but it did not by itself grant unrestricted access to a computer or repository.

Streaming was also supported, allowing partial text responses to be delivered as they were generated rather than waiting for the complete response. Structured outputs were supported for applications that needed responses conforming to a defined schema. The supplied specifications distinguish structured outputs from a separate JSON-mode field, which was not verified here.

Pricing when the model was available

Before API retirement, GPT-5.1-Codex-Max was priced at $1.25 per million input tokens, $0.125 per million cached input tokens, and $10 per million output tokens. Cached input pricing applied when eligible input could be reused according to OpenAI’s caching rules.

The pricing structure favored applications that could limit output volume and reuse stable context. Output tokens were substantially more expensive than uncached input tokens, which matters for long coding sessions that generate large patches, explanations, or repeated tool-related responses. A large context window also made it possible to send more project information, but it did not remove the need to control irrelevant context and unnecessary output.

Main strengths and trade-offs

GPT-5.1-Codex-Max’s main strength was its fit for long-horizon engineering tasks. The combination of a 400,000-token context window, 128,000-token output limit, reasoning, function calling, streaming, structured outputs, and context compaction addressed workflows that are difficult to handle with a short-context coding assistant.

It was especially well suited to tasks where the agent needed to maintain a broad understanding of a repository while making changes over many iterations. Examples include migrating patterns across many files, tracing a difficult defect through a large application, implementing a feature that touches frontend and backend code, or preparing a pull request after tests and review feedback.

The trade-off was cost and latency relative to smaller or simpler coding options. GPT-5.1-Codex-Max was intended for difficult, extended work, not necessarily for every autocomplete request or short coding question. Its output price was $10 per million tokens, and its reasoning-oriented, long-running workflow could be unnecessary when a task required only a brief answer or a small local edit. The supplied research gives comparative editorial scores of 9 for reasoning, 10 for coding, 8 for speed, and 8 for cost, but these are estimates rather than OpenAI-published ratings.

When to choose this model

Historically, GPT-5.1-Codex-Max was the better fit when the task required sustained repository awareness and multiple rounds of tool use. Appropriate examples included:

  • Refactoring or migrating a large codebase across many files
  • Debugging problems that required inspecting substantial project context
  • Implementing a feature through repeated planning, coding, testing, and revision
  • Reviewing or preparing a pull request with changes spanning several components
  • Working on frontend tasks where source code and visual input both mattered
  • Running an extended Codex session that could exceed one ordinary context window

For a short code explanation, a small edit, or a high-volume low-complexity task, a faster or less expensive coding model would generally be more appropriate. For new production integrations after July 23, 2026, another current model should be selected instead; OpenAI’s deprecation documentation listed GPT-5.6 Sol as the recommended replacement.

Limitations and current status

GPT-5.1-Codex-Max returned text only. It supported image input but did not support audio or video input, and it did not natively generate images, audio, or video. Its knowledge cutoff was September 30, 2024, so current facts or newly changed project information had to be supplied through the application, tools, or other external context.

The model was also not intended to be a general-purpose conversational model. Its design emphasized software engineering and extended agent workflows. Even with a very large context window, results still depended on the quality of the repository context, tool integration, instructions, testing, and review supplied by the surrounding application.

Most importantly, the model is retired for API use. OpenAI announced deprecation on April 22, 2026, and API access ended on July 23, 2026. It remains useful as a reference point for understanding OpenAI’s long-running coding-model design, but developers should not plan new API deployments around it.


Answers to Frequently Asked Questions

Is GPT-5.1-Codex-Max still available?
No. OpenAI announced GPT-5.1-Codex-Max’s deprecation on April 22, 2026, and API access ended on July 23, 2026. OpenAI’s deprecation documentation identified GPT-5.6 Sol as the recommended replacement, so developers should use a current model for new production integrations.
How much did GPT-5.1-Codex-Max cost?
Before API retirement, GPT-5.1-Codex-Max cost $1.25 per million input tokens, $0.125 per million cached input tokens, and $10 per million output tokens. Its pricing made it more suitable for complex, extended coding tasks than for simple autocomplete or short coding questions.
What API features did GPT-5.1-Codex-Max support?
When available, GPT-5.1-Codex-Max was accessed through OpenAI’s Responses API. It supported text and image input, text output, reasoning, function calling, streaming, and structured outputs. It did not natively generate images, audio, or video.
How large was GPT-5.1-Codex-Max’s context window?
GPT-5.1-Codex-Max had a 400,000-token context window and supported context compaction. This allowed coding agents to preserve the most relevant information and continue complex tasks across multiple context windows, potentially spanning millions of tokens.
What was GPT-5.1-Codex-Max designed for?
GPT-5.1-Codex-Max was a specialized OpenAI model for agentic coding and long-running software-engineering tasks. It was designed to inspect repositories, edit multiple files, use tools, debug issues, review changes, create pull requests, and continue working across multiple iterations.


Sources 5
Provider

About OpenAI