What was codex-mini-latest?
codex-mini-latest was OpenAI’s specialized coding model for Codex CLI and related software-engineering workflows. OpenAI described it as a fine-tuned version of o4-mini, adapted for code question answering, code editing, and fast development tasks rather than general-purpose use across every modality.
The model launched on May 16, 2025. It accepted text and image input and produced text output, allowing developers to provide source code, instructions, repository context, and supported visual references. Its main practical role was helping a coding agent inspect a project, reason about a change, propose edits, and return instructions or tool calls for work in a developer-controlled environment.
There is an important status limitation: OpenAI deprecated codex-mini-latest on November 17, 2025 and removed it from the API on February 12, 2026. OpenAI listed gpt-5-codex-mini as its replacement. Consequently, the specifications and prices below describe the retired model’s documented capabilities and historical positioning, not a currently available integration target.
Key specifications and documented capabilities
| Specification | Documented value |
|---|---|
| Provider | OpenAI |
| Model family | Codex; fine-tuned from o4-mini |
| Release date | May 16, 2025 |
| Status | Retired; API access ended February 12, 2026 |
| Context window | 200,000 tokens |
| Maximum output | 100,000 tokens |
| Input | Text and images |
| Output | Text |
| Knowledge cutoff | June 1, 2024 |
| Tools and interfaces | Function calling, streaming, batch processing, structured outputs, and local shell through the Responses API |
These are documented model specifications. They should not be confused with the editorial scores associated with this record, which rate reasoning, coding, speed, and cost comparatively and are not OpenAI benchmark results.
Coding and reasoning role
codex-mini-latest was built around software-engineering reasoning: understanding a request, examining relevant code, planning a change, and producing an implementation or explanation. Its intended tasks included code question answering, code editing, repository-level work, and other Codex CLI operations.
The model’s relationship to o4-mini is useful for positioning. OpenAI used o4-mini as the base and then specialized the resulting model for coding workflows. That specialization meant the model was not simply a generic low-cost chat model with coding among many possible uses. Its documented purpose was narrower: deliver low-latency assistance during development, where a user or coding agent may repeatedly ask for analysis, edits, and commands.
The supplied editorial assessment rates its reasoning at 7 out of 10 and coding at 8 out of 10. Those figures are comparative editorial estimates, not provider-published scores. They indicate the expected balance: capable enough for practical coding assistance, but optimized more for speed and workflow efficiency than for being the broadest or most capable model available.
Context window and output limits
The 200,000-token context window was one of the model’s most useful technical properties. A context window is the amount of text and other supported input the model can consider in one request. For coding workflows, a large window can reduce the need to split a task into many small prompts and can make it easier to provide multiple files, repository instructions, error messages, and relevant history together.
Its maximum output was documented as 100,000 tokens. That is an upper limit, not a promise that every response would be that long or that such a length is desirable. Normal code edits and explanations would generally be much shorter. The limit was most relevant to long-running coding tasks, large patches, or agent workflows that needed extensive output.
The model’s listed knowledge cutoff was June 1, 2024. Local shell execution, external repository content, and other supplied context could give a workflow access to newer information, but they did not change the underlying model knowledge cutoff.
Tools, function calling, and local shell
codex-mini-latest supported function calling, streaming, structured outputs, and batch processing. Function calling allows an application to expose defined operations that the model can request, while the application remains responsible for deciding whether and how to execute them. Streaming delivers generated output progressively rather than waiting for the complete response, which is useful for interactive coding interfaces.
Structured outputs were also supported. This capability helps an application request results that follow a defined data structure instead of relying only on free-form text. The supplied research does not establish that the model had a separate, independently documented “JSON mode,” so structured outputs should not automatically be treated as evidence of a distinct JSON-mode feature.
A particularly relevant feature was OpenAI’s local shell tool through the Responses API. With this workflow, the model could return shell commands for execution in a developer-controlled local environment. The model did not independently receive unrestricted control of the user’s computer: the surrounding application or developer-controlled environment had to handle execution. This made local-shell integration useful for repository inspection, running tests, checking files, and carrying out approved development operations while retaining an execution boundary outside the model itself.
Supported modalities and important limitations
The model supported text input and image input. Images could be useful when a developer needed to provide a visual reference, such as an interface screenshot or another image-based coding context supported by the workflow. It did not support audio or video input, and it did not generate images, audio, video, music, or speech. Its direct output was text.
It was therefore not an appropriate choice for audio assistants, video understanding, image generation, or speech applications. It was also not a general replacement for every OpenAI model. OpenAI’s documentation recommended gpt-4.1 for direct general API use at the time of the model’s documentation, while codex-mini-latest was aimed specifically at fast coding work and Codex CLI.
Historical pricing and cost trade-offs
Before retirement, the documented standard input price was $1.50 per 1 million input tokens. Cached input was priced at $0.375 per 1 million tokens, and output was priced at $6.00 per 1 million output tokens.
Cached input pricing mattered for repeated workflows that reused the same large prompt or repository context. When a coding application sent recurring instructions or stable project information, eligible cached content could cost less than ordinary input. Output remained more expensive per token than standard input, so applications could control costs by requesting focused patches and concise explanations when very long responses were unnecessary.
The supplied record describes the model as relatively fast and cost-efficient for its coding role. Its editorial scores are 8 out of 10 for speed and 7 out of 10 for cost. Again, these are subjective comparative evaluations rather than OpenAI’s official performance measurements. The practical trade-off was a lower-latency coding specialist versus a broader or more capable model that might be preferable for especially difficult reasoning tasks.
Best use cases
- Codex CLI workflows: interactive code questions, edits, repository navigation, and development assistance from a command-line environment.
- Repository tasks: analyzing relevant project files, proposing changes, and helping developers work through implementation tasks.
- Fast code editing: situations where response speed and repeated interaction matter more than maximum general reasoning capability.
- Agent-style development: applications that combine model responses with function calls, structured results, streaming, batch processing, or controlled local-shell execution.
- Large code context: tasks that benefit from a documented 200,000-token context window, such as considering multiple files or substantial project instructions together.
When would it have been the right choice?
When it was available, codex-mini-latest was a sensible choice for developers who wanted a coding-focused model with fast interaction, a large context window, and Codex CLI-oriented tooling. It was especially well suited to code Q&A, routine editing, repository exploration, and workflows where many small model interactions were more valuable than a single extremely deep response.
A broader general-purpose model was more appropriate for applications not centered on software engineering, and a model with audio or video capabilities was necessary for those modalities. Developers needing a currently available API model should not choose codex-mini-latest at all, because its API access ended on February 12, 2026. For a successor within the documented product direction, OpenAI identified gpt-5-codex-mini, although the supplied research does not provide that model’s specifications or pricing and therefore does not support a detailed comparison.
Bottom line
codex-mini-latest was a focused, fast coding model rather than a general catalog entry. Its strongest documented advantages were its Codex CLI positioning, 200,000-token context window, 100,000-token maximum output, text-and-image input, and support for coding-agent tools including controlled local-shell workflows. Its limitations included text-only output, no audio or video support, a coding-specific scope, and a June 1, 2024 knowledge cutoff. Most importantly for practical planning, the model is retired: its historical capabilities may explain its role, but new integrations need a currently supported replacement.

