What GPT-5.1-Codex Mini was
GPT-5.1-Codex Mini was an OpenAI coding model and a smaller member of the GPT-5.1-Codex family. OpenAI released it on November 13, 2025. Its canonical API identifier was gpt-5.1-codex-mini.
The model was intended for agentic coding: software tasks in which the model does more than answer a single programming question. In a Codex-style workflow, it can reason about a repository, inspect or modify code through tools, keep track of a long sequence of actions, and return an implementation or explanation. Typical tasks included code editing, repository maintenance, refactoring, bug fixing, and other repeated coding operations.
“Mini” described its position within the family rather than a lack of practical utility. Compared with a larger coding model, it was designed to exchange some capability for lower cost and higher speed. That trade-off made it suitable for workloads involving many model calls or long-running automated agents, provided the task did not require the strongest available coding performance.
Lifecycle and availability
GPT-5.1-Codex Mini is retired. The supplied lifecycle data lists April 22, 2026 as its deprecation date and July 23, 2026 as the date API access shut down. OpenAI recommended GPT-5.6 Terra as a replacement in the supplied documentation.
This status is important for practical evaluation. The model’s historical specifications and pricing remain useful for understanding its design, but its low price and technical capabilities should not be treated as an invitation to start a new deployment. New applications should use an available replacement after checking the replacement’s own documentation, pricing, limits, and coding behavior.
Context window, output limit, and knowledge cutoff
| Specification | Reported value |
|---|---|
| Context length | 400,000 tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | September 30, 2024 |
| Model type | Coding |
| Release date | November 13, 2025 |
| Provider | OpenAI |
The 400,000-token context window was one of the model’s most useful properties for repository work. A context window is the amount of information the model can consider in one request, including instructions, conversation history, source files, tool results, and other supplied content. A large window can reduce the need to summarize or split a complex codebase into many small requests, although it does not guarantee that every detail will receive equal attention.
The maximum output limit was 128,000 tokens. This gave an agent room to produce substantial patches, explanations, or intermediate reasoning-related output where applicable, but applications still needed to manage output length and validate changes rather than assuming that a long response was correct.
OpenAI listed a September 30, 2024 knowledge cutoff. That cutoff describes the model’s stored knowledge and does not mean the model could automatically browse the web or know later events. The supplied research does not verify web-search support, so web access should not be assumed.
Pricing and cost position
The reported historical API price was $0.25 per 1 million input tokens and $2.00 per 1 million output tokens. Cached input tokens were priced at $0.025 per 1 million tokens. These were token-based API prices rather than a subscription plan or a per-user monthly fee.
| Token type | Historical price |
|---|---|
| Input | $0.25 per 1 million tokens |
| Cached input | $0.025 per 1 million tokens |
| Output | $2.00 per 1 million tokens |
The pricing structure favored workloads that reused substantial prompt material, such as a repository context or stable system instructions. Prompt caching could reduce the cost of repeated input, while batch processing could help with suitable non-interactive workloads. The model’s low reported input price and relatively fast response profile made it attractive for high-volume coding automation, but those historical prices should not be used to estimate a currently available service.
Supported inputs, outputs, and tools
GPT-5.1-Codex Mini accepted text and images. Image input could be useful when a coding task involved a screenshot, diagram, visual test failure, or interface reference. Audio and video inputs were not supported according to the supplied research.
The model generated text only. It did not natively generate images, audio, or video, so it was not an appropriate choice for media-generation workflows. Its multimodal capability was therefore asymmetric: it could inspect an image alongside a coding request, but its response remained text such as code, explanations, plans, or tool-call instructions.
GPT-5.1-Codex Mini supported tool use and function calling. In practical terms, a surrounding application could expose operations such as reading files, searching a repository, applying edits, running tests, or interacting with other developer tools. The model’s ability to call a function did not by itself make those operations safe or available; the application had to define the tools, enforce permissions, execute calls, and validate results.
Streaming was supported, allowing an application to receive generated output incrementally instead of waiting for the entire response. The model also supported structured outputs, which could help an application request responses that follow a specified schema. The supplied data does not verify a separate JSON mode, so structured outputs should not automatically be described as a distinct JSON-mode feature.
Reasoning and coding performance
The model supported reasoning tokens, meaning it could spend part of its generation process working through a problem before presenting its answer or action. This was relevant to multi-step coding tasks where the model needed to inspect dependencies, identify likely causes, choose an edit strategy, and respond to test results.
GPT-5.1-Codex Mini’s primary capability was coding rather than general media generation. It was aimed at repository-level work, code editing, and maintenance tasks that benefit from repeated interaction with tools. Its large context and agent-oriented design could be more useful than a short-answer model when the task required examining many files or preserving a long chain of context.
The supplied editorial assessment gave it a reasoning score of 7 out of 10, a coding score of 8 out of 10, and a speed score of 9 out of 10. These are comparative editorial estimates, not OpenAI-published benchmark results. The same assessment gave it a cost score of 9 out of 10. They should be read as a summary of expected trade-offs: fast and inexpensive for its category, with solid coding ability, but not the strongest option for the hardest software-engineering problems.
Main strengths and limitations
Where it was strongest
- Agentic coding: It was built for workflows in which an agent repeatedly examines code, calls tools, edits files, and responds to test or execution results.
- Long tasks: The 400,000-token context window could accommodate substantial project context, long conversations, and many tool results.
- Cost-sensitive automation: Its reported input and output prices were positioned for applications making many coding requests.
- Fast iteration: The editorial speed estimate rated it highly, making it a reasonable fit for interactive code assistance and frequent agent steps.
- Operational features: Streaming, function calling, structured outputs, caching, and batch processing supported production-style orchestration.
- Image-aware coding: It could accept image input when a developer needed to provide a screenshot or visual reference with a programming task.
Where it was limited
- Retirement: API access ended on July 23, 2026 according to the supplied lifecycle information, making it unsuitable for new deployments.
- Capability ceiling: As the smaller GPT-5.1-Codex model, it traded some capability for speed and cost. More difficult coding or reasoning tasks could require a stronger model.
- Text-only output: It did not produce images, audio, or video.
- No audio or video input: Media-heavy workflows were outside its supported input modalities.
- No fine-tuning: The supplied specification marks fine-tuning as unsupported.
- Knowledge cutoff: Its listed knowledge cutoff was September 30, 2024, so later information had to be supplied through prompts, tools, or other application logic.
When to choose GPT-5.1-Codex Mini
Historically, GPT-5.1-Codex Mini was a good fit when the main objective was to complete coding work quickly and economically rather than maximize peak reasoning or code-generation quality. Examples included automated repository cleanup, routine code transformations, maintenance tickets, test-oriented edits, and agent workflows that made many model calls over a large project.
It was especially appropriate when prompt caching could be used effectively, when the application benefited from streaming responses, or when a developer needed function calls and structured results for an automated coding harness. Image input also made it useful for tasks that combined a visual bug report with source-code changes.
A larger or newer coding model was more appropriate when the task involved unusually difficult architecture decisions, subtle debugging, high-risk production changes, or a need for the strongest available reasoning and coding performance. A general multimodal model would be more suitable for workflows requiring audio or video input, while a media-generation model would be needed for image, audio, or video output.
Because GPT-5.1-Codex Mini is retired, the most important present-day choice is not whether its historical price or speed trade-off is attractive. It is whether the recommended replacement provides the required context size, coding quality, tool support, modality coverage, latency, and cost for the application. Existing users should migrate rather than build new dependencies on the retired model.
Bottom line
GPT-5.1-Codex Mini was a focused engineering model: fast, relatively inexpensive, and designed to keep long-running coding agents moving through repositories and tool calls. Its 400,000-token context, 128,000-token output limit, image input, reasoning support, and function calling made it more suitable for agentic software work than for ordinary short programming questions alone. Its limitations were equally clear: it was less capable than larger Codex options, produced text only, lacked audio and video support, and is now retired. The model is best understood as a historical example of the speed-and-cost end of OpenAI’s Codex lineup, not as a current production choice.

