What GPT-5-Codex was built to do
GPT-5-Codex was a GPT-5-family model from OpenAI, optimized specifically for agentic software engineering. In this context, “agentic” means that the model was intended to participate in a multi-step task rather than simply answer a single coding question. A typical workflow could involve understanding a repository, identifying the relevant files, proposing or applying a change, running or interpreting tests through tools, and revising the implementation.
OpenAI positioned GPT-5-Codex for both interactive and autonomous coding work. It could assist with a quick edit or code review in a developer-facing environment, but it was also aimed at longer tasks where the model had to maintain context and make progress with less continuous human direction. The model was made available through Codex surfaces and later through the Responses API.
This specialization is important for understanding the product. GPT-5-Codex was not primarily a general-purpose chatbot, image generator, or audio model. Its defining use was software engineering carried out with reasoning, repository context, and tool support.
Status and position in OpenAI’s catalog
GPT-5-Codex was announced on September 15, 2025. OpenAI made it available through the Responses API on September 23, 2025, using the canonical model identifier gpt-5-codex. The model was subsequently deprecated and its API access ended on July 23, 2026.
As a result, GPT-5-Codex no longer belongs in a new production architecture. Its former role was that of a coding-specialized GPT-5 option for Codex and similar development environments. Developers evaluating current models should select an actively supported coding or general reasoning model instead of attempting to build around this identifier.
Technical specifications
The following are the documented characteristics of GPT-5-Codex before retirement:
| Specification | Documented value |
|---|---|
| Provider | OpenAI |
| Model family | GPT-5 |
| Primary type | Coding and software engineering |
| Model ID | gpt-5-codex |
| Context window | 400,000 tokens |
| Maximum output | 128,000 tokens |
| Input | Text and images |
| Output | Text |
| Knowledge cutoff | September 30, 2024 |
| Fine-tuning | Not supported |
| API availability | Responses API before shutdown |
A token is a unit of text used for context and billing; it can be a word, part of a word, punctuation, or another small text segment. The 400,000-token context window was large enough for substantial repository material and extended task history, although the practical amount a developer could provide still depended on the surrounding instructions, tool results, and requested output.
The 128,000-token maximum output was a ceiling, not a recommendation that every response should be that long. Most code changes, explanations, and review results would use far less. The limit was most relevant to long-running tasks that needed to return extensive analysis, patches, or generated code.
Coding, reasoning, and tool capabilities
GPT-5-Codex supported reasoning-token processing, allowing it to spend internal computation working through complex problems before producing an answer. For software engineering, that capability was relevant to tasks such as tracing a bug across multiple files, planning a refactor, understanding dependencies, or deciding how a change could affect existing tests.
Its coding focus covered repository-level engineering rather than only isolated snippets. Appropriate tasks included reviewing a pull request, debugging a failing implementation, restructuring code, generating tests, and implementing a frontend change from a screenshot or design reference. These workflows benefit from a model that can connect a requested change to the broader structure of an application.
The model also supported function calling and tool use. Function calling allows an application to give the model access to defined operations, such as repository inspection, test execution, or other development tools. The model could decide when one of those operations was useful and incorporate the returned information into its next step. The available tools themselves depended on the surrounding Codex or API integration; the model did not automatically provide unrestricted access to a developer’s computer or repository.
Additional documented capabilities included streaming, structured outputs, and batch processing. Streaming allowed partial text to be delivered while a response was being generated. Structured outputs helped applications request responses that followed a defined schema. Batch processing was useful for submitting groups of jobs, although the supplied documentation does not establish that GPT-5-Codex had a separate, independently verified JSON mode. It was also not available for fine-tuning.
Supported input and output modalities
GPT-5-Codex accepted text and image input and returned text output. Image support made it useful for development tasks involving screenshots, frontend designs, visual bugs, or user-interface references. For example, a developer could use a screenshot as a reference while asking the model to implement or revise a page.
The model did not support audio or video input, and it did not generate images, audio, or video. It should therefore not be selected for multimedia production workflows or for applications that require speech or video understanding. Its multimodal capability was specifically useful because visual input could complement a coding task; it did not turn GPT-5-Codex into a general visual-generation system.
Pricing before retirement
Before the model was shut down, OpenAI listed standard usage pricing of $1.25 per million input tokens, $0.125 per million cached input tokens, and $10 per million output tokens.
| Usage type | Price before shutdown |
|---|---|
| Input tokens | $1.25 per 1 million tokens |
| Cached input tokens | $0.125 per 1 million tokens |
| Output tokens | $10.00 per 1 million tokens |
These were token-based API prices, not a subscription fee or a current purchasing option. Output tokens cost substantially more than ordinary input tokens, so long generated responses could have a noticeable effect on total usage. Caching could reduce the price of repeated input when the relevant content qualified for cached-input billing. Because API access ended on July 23, 2026, these figures are historical and cannot be used to plan a new deployment.
Main strengths and limitations
GPT-5-Codex’s main strength was the combination of coding specialization, reasoning, tool use, and a large context window. That combination suited work where the model had to understand more than one function or file. Repository-level debugging, broad refactoring, test generation, and code review could all require maintaining relationships across a sizable amount of project material.
Visual input was another practical advantage for frontend work. A screenshot or design reference could provide information that is difficult to express completely in text, such as spacing, layout, component appearance, or a visible rendering defect.
There were also clear limitations. The model’s knowledge cutoff was September 30, 2024, so its built-in knowledge was not a reliable source for developments after that date. Tool access could provide current project information, but it did not change the model’s underlying training cutoff. GPT-5-Codex also lacked audio and video input, could not generate images, and did not support fine-tuning.
Most importantly, retirement overrides its former technical strengths. A model can be well suited to a task in principle and still be unsuitable for production if its endpoint is unavailable. The shutdown date should therefore be treated as a decisive limitation rather than a minor lifecycle note.
Best use cases before shutdown
- Repository-level coding: understanding and modifying related files across a software project.
- Code review: examining proposed changes, identifying likely defects, and explaining maintainability concerns.
- Debugging: tracing failures through code and using tool results or tests to refine a diagnosis.
- Refactoring: restructuring code while considering interactions between modules and existing behavior.
- Test generation: creating tests for new or changed functionality.
- Frontend implementation: translating screenshots or visual references into code and revising the result based on visual requirements.
- Longer autonomous engineering tasks: carrying out multiple related steps with limited supervision through a Codex-style workflow.
These use cases describe the model’s intended role before retirement. They do not imply that the former endpoint remains available.
When another option is more appropriate
For a current deployment, an actively maintained coding model is more appropriate than GPT-5-Codex because the latter can no longer be accessed. This is true even when a replacement has a different context limit or pricing structure.
More generally, a coding-specialized agentic model is most useful when the task involves multiple files, tool calls, sustained reasoning, or a large amount of project context. A smaller or faster coding model may be preferable for simple completions, short edits, or high-volume requests where latency and cost matter more than extended planning. A general-purpose model may be a better fit when the application combines coding with requirements outside GPT-5-Codex’s supported modalities, such as audio or video understanding.
For current systems, the important comparison is therefore not just raw capability. Developers should weigh active availability, coding performance, context requirements, tool integration, response speed, and token cost. GPT-5-Codex was designed around the capability side of that trade-off, but its retirement makes lifecycle support the overriding consideration.
Bottom line
GPT-5-Codex was OpenAI’s purpose-built GPT-5 model for agentic software engineering. It supported text and image input, text output, reasoning, function calling, streaming, structured outputs, batch processing, a 400,000-token context window, and up to 128,000 output tokens. Its design made it suitable for repository work, debugging, code review, refactoring, testing, and visual frontend tasks. However, the model’s API access ended on July 23, 2026, so it is now relevant as a documented and historical model rather than as a viable choice for new software.

