What GPT-5.1-Codex-Max was designed to do
GPT-5.1-Codex-Max was a specialized OpenAI model for agentic coding. In this context, agentic coding means the model can work through a software task in multiple stages: inspect a repository, reason about the required changes, edit several files, use tools, review the result, and continue based on feedback.
Rather than focusing primarily on short conversational answers or isolated code completion, GPT-5.1-Codex-Max was optimized for sustained software-engineering work inside Codex. Its target tasks included multi-file implementation, repository-scale refactoring, deep debugging, code review, frontend development, pull-request creation, and technical question answering.
The model launched in Codex on November 19, 2025. It became available through OpenAI’s API in December 2025, using the Responses API. OpenAI later announced its deprecation on April 22, 2026, and API access ended on July 23, 2026. The deprecation documentation identified GPT-5.6 Sol as the recommended replacement, so GPT-5.1-Codex-Max should now be treated as a retired model rather than a new production option.
Long-running context and compaction
The model’s defining capability was support for work across multiple context windows. A context window is the amount of conversation, code, tool output, and other information the model can consider in one request. GPT-5.1-Codex-Max exposed a 400,000-token context window and supported a process OpenAI called compaction.
Compaction preserves the most relevant information from an earlier part of an agent session so work can continue in a later context window. This is important for large repositories and debugging sessions where the complete task history may eventually exceed the limit of one request. OpenAI described the model as capable of coherent work spanning millions of tokens across a single task through this multi-window workflow.
Compaction did not mean that every original token remained available in full detail forever. As with any context-management process, the agent still depended on retaining the information most relevant to the next stage. The practical advantage was that a long task could continue without restarting from scratch whenever one individual context window became full.
Coding and reasoning capabilities
GPT-5.1-Codex-Max was trained on software-engineering tasks as well as mathematics, research, computer use, and related reasoning domains. Its coding orientation made it particularly relevant when the desired result was a working change across a codebase rather than a short explanatory response.
- Repository-scale refactoring and migration work
- Multi-file implementation and debugging
- Code review and pull-request creation
- Frontend and application development
- Extended agent loops using tools and iterative feedback
- Technical question answering connected to software projects
Reasoning tokens were supported, allowing the model to spend additional computation on problems that required planning or analysis. The supplied research does not provide a single benchmark score or a guaranteed completion rate, so claims about reasoning quality should be understood as positioning and capability descriptions rather than a promise of a specific outcome.
API specifications and supported inputs
When it was available, GPT-5.1-Codex-Max was accessed through the Responses API. It accepted text and images as input and returned text. Image input could therefore provide visual context alongside code or instructions, but the model was not a native image, audio, or video generation system.
| Specification | Value |
|---|---|
| Context window | 400,000 tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | September 30, 2024 |
| Input modalities | Text and image |
| Output modality | Text |
| API | Responses API |
| Reasoning | Supported |
| Function calling | Supported |
| Streaming | Supported |
| Structured outputs | Supported |
The 128,000-token maximum output applied to an individual model context. It should not be interpreted as a guarantee that every request could produce that much useful code or text. Output length also depended on the task, instructions, available context, and the agent’s tool workflow.
Tools, structured output, and agent workflows
Function calling allowed an application to expose tools that the model could request during a task. In a coding workflow, those tools could support actions such as inspecting files, running checks, or interacting with an external development system, subject to the application’s implementation and permissions. The model’s support for function calling was therefore useful for building an agent loop, but it did not by itself grant unrestricted access to a computer or repository.
Streaming was also supported, allowing partial text responses to be delivered as they were generated rather than waiting for the complete response. Structured outputs were supported for applications that needed responses conforming to a defined schema. The supplied specifications distinguish structured outputs from a separate JSON-mode field, which was not verified here.
Pricing when the model was available
Before API retirement, GPT-5.1-Codex-Max was priced at $1.25 per million input tokens, $0.125 per million cached input tokens, and $10 per million output tokens. Cached input pricing applied when eligible input could be reused according to OpenAI’s caching rules.
The pricing structure favored applications that could limit output volume and reuse stable context. Output tokens were substantially more expensive than uncached input tokens, which matters for long coding sessions that generate large patches, explanations, or repeated tool-related responses. A large context window also made it possible to send more project information, but it did not remove the need to control irrelevant context and unnecessary output.
Main strengths and trade-offs
GPT-5.1-Codex-Max’s main strength was its fit for long-horizon engineering tasks. The combination of a 400,000-token context window, 128,000-token output limit, reasoning, function calling, streaming, structured outputs, and context compaction addressed workflows that are difficult to handle with a short-context coding assistant.
It was especially well suited to tasks where the agent needed to maintain a broad understanding of a repository while making changes over many iterations. Examples include migrating patterns across many files, tracing a difficult defect through a large application, implementing a feature that touches frontend and backend code, or preparing a pull request after tests and review feedback.
The trade-off was cost and latency relative to smaller or simpler coding options. GPT-5.1-Codex-Max was intended for difficult, extended work, not necessarily for every autocomplete request or short coding question. Its output price was $10 per million tokens, and its reasoning-oriented, long-running workflow could be unnecessary when a task required only a brief answer or a small local edit. The supplied research gives comparative editorial scores of 9 for reasoning, 10 for coding, 8 for speed, and 8 for cost, but these are estimates rather than OpenAI-published ratings.
When to choose this model
Historically, GPT-5.1-Codex-Max was the better fit when the task required sustained repository awareness and multiple rounds of tool use. Appropriate examples included:
- Refactoring or migrating a large codebase across many files
- Debugging problems that required inspecting substantial project context
- Implementing a feature through repeated planning, coding, testing, and revision
- Reviewing or preparing a pull request with changes spanning several components
- Working on frontend tasks where source code and visual input both mattered
- Running an extended Codex session that could exceed one ordinary context window
For a short code explanation, a small edit, or a high-volume low-complexity task, a faster or less expensive coding model would generally be more appropriate. For new production integrations after July 23, 2026, another current model should be selected instead; OpenAI’s deprecation documentation listed GPT-5.6 Sol as the recommended replacement.
Limitations and current status
GPT-5.1-Codex-Max returned text only. It supported image input but did not support audio or video input, and it did not natively generate images, audio, or video. Its knowledge cutoff was September 30, 2024, so current facts or newly changed project information had to be supplied through the application, tools, or other external context.
The model was also not intended to be a general-purpose conversational model. Its design emphasized software engineering and extended agent workflows. Even with a very large context window, results still depended on the quality of the repository context, tool integration, instructions, testing, and review supplied by the surrounding application.
Most importantly, the model is retired for API use. OpenAI announced deprecation on April 22, 2026, and API access ended on July 23, 2026. It remains useful as a reference point for understanding OpenAI’s long-running coding-model design, but developers should not plan new API deployments around it.

