Codex

codex-mini-latest

by OpenAI · Retired; API access ended on 2026-02-12

codex-mini-latest was OpenAI’s fast, coding-focused reasoning model for Codex CLI, repository tasks, and code editing. It supported a 200,000-token context, up to 100,000 output tokens, image input, function calling, structured outputs, streaming, batch processing, and controlled local-shell workflows. It was deprecated in November 2025 and removed from the API on February 12, 2026.

Text Reasoning Coding
codex-mini-latest was a specialized OpenAI coding model designed for fast code question answering, editing, repository tasks, and Codex CLI use. Its large context window and local-shell integration made it suitable for agent-style development workflows, while its lower cost and latency positioned it below larger, more general-purpose coding systems. The model is now retired and should not be selected for new API integrations.
Outputs

What codex-mini-latest can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

7/10 Reasoning
8/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Codex
Model type Coding
Context window 200K tokens
Maximum output 100K tokens
Knowledge cutoff 2024-06-01
Release date 2025-05-16
Status Retired; API access ended on 2026-02-12
Deprecation date 2025-11-17
Shutdown date 2026-02-12
Knowledge cutoff notes

The official model documentation lists June 1, 2024 as the knowledge cutoff. This cutoff describes the underlying model and is not changed by tools, local shell execution, or external context supplied during use.

Model notes

codex-mini-latest was a fine-tuned version of o4-mini designed specifically for Codex CLI. OpenAI described it as optimized for fast coding workflows and low-latency code Q&A and editing. It supported text input and output plus image input, but not audio or video. The model supported streaming, function calling, structured outputs, and batch processing. OpenAI's local shell tool was available through the Responses API specifically with this model and enabled the model to return shell commands for execution in a developer-controlled local environment. OpenAI recommended gpt-4.1 for direct general API use at the time of the model documentation. The model was deprecated on November 17, 2025 and removed from the API on February 12, 2026; OpenAI listed gpt-5-codex-mini as the replacement. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input $1.50 per 1M input tokens; $0.375 per 1M cached input tokens
Output $6.00 per 1M output tokens
Model guide

codex-mini-latest: OpenAI’s Fast Coding Model for Codex CLI

codex-mini-latest was OpenAI’s fast reasoning model for Codex CLI and software-engineering workflows. Fine-tuned from o4-mini, it supported text and image input, a 200,000-token context window, up to 100,000 output tokens, function calling, structured outputs, streaming, batch processing, and local-shell workflows. It launched on May 16, 2025, was deprecated on November 17, 2025, and was removed from the API on February 12, 2026.

What was codex-mini-latest?

codex-mini-latest was OpenAI’s specialized coding model for Codex CLI and related software-engineering workflows. OpenAI described it as a fine-tuned version of o4-mini, adapted for code question answering, code editing, and fast development tasks rather than general-purpose use across every modality.

The model launched on May 16, 2025. It accepted text and image input and produced text output, allowing developers to provide source code, instructions, repository context, and supported visual references. Its main practical role was helping a coding agent inspect a project, reason about a change, propose edits, and return instructions or tool calls for work in a developer-controlled environment.

There is an important status limitation: OpenAI deprecated codex-mini-latest on November 17, 2025 and removed it from the API on February 12, 2026. OpenAI listed gpt-5-codex-mini as its replacement. Consequently, the specifications and prices below describe the retired model’s documented capabilities and historical positioning, not a currently available integration target.

Key specifications and documented capabilities

SpecificationDocumented value
ProviderOpenAI
Model familyCodex; fine-tuned from o4-mini
Release dateMay 16, 2025
StatusRetired; API access ended February 12, 2026
Context window200,000 tokens
Maximum output100,000 tokens
InputText and images
OutputText
Knowledge cutoffJune 1, 2024
Tools and interfacesFunction calling, streaming, batch processing, structured outputs, and local shell through the Responses API

These are documented model specifications. They should not be confused with the editorial scores associated with this record, which rate reasoning, coding, speed, and cost comparatively and are not OpenAI benchmark results.

Coding and reasoning role

codex-mini-latest was built around software-engineering reasoning: understanding a request, examining relevant code, planning a change, and producing an implementation or explanation. Its intended tasks included code question answering, code editing, repository-level work, and other Codex CLI operations.

The model’s relationship to o4-mini is useful for positioning. OpenAI used o4-mini as the base and then specialized the resulting model for coding workflows. That specialization meant the model was not simply a generic low-cost chat model with coding among many possible uses. Its documented purpose was narrower: deliver low-latency assistance during development, where a user or coding agent may repeatedly ask for analysis, edits, and commands.

The supplied editorial assessment rates its reasoning at 7 out of 10 and coding at 8 out of 10. Those figures are comparative editorial estimates, not provider-published scores. They indicate the expected balance: capable enough for practical coding assistance, but optimized more for speed and workflow efficiency than for being the broadest or most capable model available.

Context window and output limits

The 200,000-token context window was one of the model’s most useful technical properties. A context window is the amount of text and other supported input the model can consider in one request. For coding workflows, a large window can reduce the need to split a task into many small prompts and can make it easier to provide multiple files, repository instructions, error messages, and relevant history together.

Its maximum output was documented as 100,000 tokens. That is an upper limit, not a promise that every response would be that long or that such a length is desirable. Normal code edits and explanations would generally be much shorter. The limit was most relevant to long-running coding tasks, large patches, or agent workflows that needed extensive output.

The model’s listed knowledge cutoff was June 1, 2024. Local shell execution, external repository content, and other supplied context could give a workflow access to newer information, but they did not change the underlying model knowledge cutoff.

Tools, function calling, and local shell

codex-mini-latest supported function calling, streaming, structured outputs, and batch processing. Function calling allows an application to expose defined operations that the model can request, while the application remains responsible for deciding whether and how to execute them. Streaming delivers generated output progressively rather than waiting for the complete response, which is useful for interactive coding interfaces.

Structured outputs were also supported. This capability helps an application request results that follow a defined data structure instead of relying only on free-form text. The supplied research does not establish that the model had a separate, independently documented “JSON mode,” so structured outputs should not automatically be treated as evidence of a distinct JSON-mode feature.

A particularly relevant feature was OpenAI’s local shell tool through the Responses API. With this workflow, the model could return shell commands for execution in a developer-controlled local environment. The model did not independently receive unrestricted control of the user’s computer: the surrounding application or developer-controlled environment had to handle execution. This made local-shell integration useful for repository inspection, running tests, checking files, and carrying out approved development operations while retaining an execution boundary outside the model itself.

Supported modalities and important limitations

The model supported text input and image input. Images could be useful when a developer needed to provide a visual reference, such as an interface screenshot or another image-based coding context supported by the workflow. It did not support audio or video input, and it did not generate images, audio, video, music, or speech. Its direct output was text.

It was therefore not an appropriate choice for audio assistants, video understanding, image generation, or speech applications. It was also not a general replacement for every OpenAI model. OpenAI’s documentation recommended gpt-4.1 for direct general API use at the time of the model’s documentation, while codex-mini-latest was aimed specifically at fast coding work and Codex CLI.

Historical pricing and cost trade-offs

Before retirement, the documented standard input price was $1.50 per 1 million input tokens. Cached input was priced at $0.375 per 1 million tokens, and output was priced at $6.00 per 1 million output tokens.

Cached input pricing mattered for repeated workflows that reused the same large prompt or repository context. When a coding application sent recurring instructions or stable project information, eligible cached content could cost less than ordinary input. Output remained more expensive per token than standard input, so applications could control costs by requesting focused patches and concise explanations when very long responses were unnecessary.

The supplied record describes the model as relatively fast and cost-efficient for its coding role. Its editorial scores are 8 out of 10 for speed and 7 out of 10 for cost. Again, these are subjective comparative evaluations rather than OpenAI’s official performance measurements. The practical trade-off was a lower-latency coding specialist versus a broader or more capable model that might be preferable for especially difficult reasoning tasks.

Best use cases

  • Codex CLI workflows: interactive code questions, edits, repository navigation, and development assistance from a command-line environment.
  • Repository tasks: analyzing relevant project files, proposing changes, and helping developers work through implementation tasks.
  • Fast code editing: situations where response speed and repeated interaction matter more than maximum general reasoning capability.
  • Agent-style development: applications that combine model responses with function calls, structured results, streaming, batch processing, or controlled local-shell execution.
  • Large code context: tasks that benefit from a documented 200,000-token context window, such as considering multiple files or substantial project instructions together.

When would it have been the right choice?

When it was available, codex-mini-latest was a sensible choice for developers who wanted a coding-focused model with fast interaction, a large context window, and Codex CLI-oriented tooling. It was especially well suited to code Q&A, routine editing, repository exploration, and workflows where many small model interactions were more valuable than a single extremely deep response.

A broader general-purpose model was more appropriate for applications not centered on software engineering, and a model with audio or video capabilities was necessary for those modalities. Developers needing a currently available API model should not choose codex-mini-latest at all, because its API access ended on February 12, 2026. For a successor within the documented product direction, OpenAI identified gpt-5-codex-mini, although the supplied research does not provide that model’s specifications or pricing and therefore does not support a detailed comparison.

Bottom line

codex-mini-latest was a focused, fast coding model rather than a general catalog entry. Its strongest documented advantages were its Codex CLI positioning, 200,000-token context window, 100,000-token maximum output, text-and-image input, and support for coding-agent tools including controlled local-shell workflows. Its limitations included text-only output, no audio or video support, a coding-specific scope, and a June 1, 2024 knowledge cutoff. Most importantly for practical planning, the model is retired: its historical capabilities may explain its role, but new integrations need a currently supported replacement.


Answers to Frequently Asked Questions

Is codex-mini-latest still available through the API?
No. OpenAI deprecated codex-mini-latest on November 17, 2025, and removed it from the API on February 12, 2026. OpenAI listed gpt-5-codex-mini as its replacement, but the available documentation does not provide detailed specifications or pricing for that successor.
What was codex-mini-latest used for?
codex-mini-latest was OpenAI’s specialized coding model for Codex CLI and software-engineering workflows. It was designed for code question answering, code editing, repository analysis, fast development tasks, and agent-style workflows using tools such as function calling and controlled local-shell execution.
What were the context window and output limits of codex-mini-latest?
codex-mini-latest had a documented context window of 200,000 tokens and a maximum output limit of 100,000 tokens. The large context window supported repository-level tasks involving multiple files, instructions, error messages, and project history.
How much did codex-mini-latest cost before it was retired?
Before retirement, standard input cost $1.50 per 1 million tokens, cached input cost $0.375 per 1 million tokens, and output cost $6.00 per 1 million tokens. These prices are historical and no longer apply because the model’s API access ended on February 12, 2026.
What tools and modalities did codex-mini-latest support?
The model accepted text and image input and produced text output. It supported function calling, streaming, batch processing, structured outputs, and local shell workflows through the Responses API. Local shell commands required execution by a developer-controlled application or environment; the model did not have unrestricted access to a user’s computer.


Sources 4
Provider

About OpenAI