o4

o4-mini

by OpenAI · Deprecated; currently available through the API; scheduled for shutdown on 2026-10-23

OpenAI o4-mini is a text-and-image reasoning model with a 200,000-token context window, up to 100,000 output tokens, coding and visual-analysis capabilities, function calling, structured outputs, streaming, batch processing, and documented reinforcement fine-tuning. It costs $1.10 per million input tokens and $4.40 per million output tokens. The model is deprecated and scheduled for API shutdown on October 23, 2026, so new deployments should include a migration plan.

Text Reasoning Coding
OpenAI o4-mini is a small o-series reasoning model for applications that need deliberate problem solving without the price or latency of a larger model. It is particularly suited to coding, mathematics, image and chart interpretation, structured extraction, and tool-using workflows. The model accepts text and images and produces text, but its API lifecycle is limited: OpenAI has deprecated it and lists October 23, 2026 as the scheduled shutdown date.
Outputs

What o4-mini can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Fine-tuning Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family o4
Model type Reasoning
Context window 200K tokens
Maximum output 100K tokens
Knowledge cutoff 2024-06-01
Release date 2025-04-16
Status Deprecated; currently available through the API; scheduled for shutdown on 2026-10-23
Shutdown date 2026-10-23
Knowledge cutoff notes

The official model page lists June 1, 2024 as the knowledge cutoff. Web search and other retrieval tools can provide newer information during use but do not change the underlying cutoff.

Model notes

OpenAI released o4-mini on April 16, 2025. The model accepts text and images and returns text. OpenAI's model page lists the unversioned o4-mini alias as deprecated and identifies GPT-5 Mini as its successor. The dated snapshot o4-mini-2025-04-16 is also deprecated. OpenAI retired o4-mini from ChatGPT on February 13, 2026, while API access continued. Current web-search documentation lists October 23, 2026 as the shutdown date. Standard pricing is separate from tool-call charges. Structured outputs are documented; a distinct legacy JSON-mode capability was not independently verified. Reinforcement fine-tuning is documented for o4-mini-2025-04-16.

Cost

Model pricing

Input $1.10 per 1 million input tokens; $0.275 per 1 million cached input tokens
Output $4.40 per 1 million output tokens
Model guide

o4-mini: Fast, Low-Cost Reasoning for Coding and Visual Analysis

OpenAI o4-mini is a compact reasoning model designed for fast, cost-efficient performance in mathematics, coding, visual reasoning, and high-volume API workloads. It accepts text and image input, returns text, supports tool calling and structured outputs, and offers a 200,000-token context window with up to 100,000 output tokens. Although it remains available through the API, it is deprecated and scheduled for shutdown on October 23, 2026.

What is OpenAI o4-mini?

o4-mini is a compact reasoning model from OpenAI, released on April 16, 2025. Unlike a conventional lightweight text-generation model, it is designed to spend additional computation working through multi-step problems before producing an answer. That makes it useful for tasks such as mathematical analysis, code generation and review, visual question answering, and structured decision-making.

The model is intended to occupy a practical middle ground: it offers stronger reasoning than many basic low-cost language models while aiming for higher speed and lower cost than larger reasoning systems. OpenAI presents it for mathematics, coding, visual reasoning, data science, and high-volume API applications where every request does not need the deepest available reasoning.

The canonical API name is o4-mini. OpenAI also lists the dated snapshot o4-mini-2025-04-16 separately. Both the unversioned model and dated snapshot are deprecated, so new production integrations should account for migration rather than treating o4-mini as a long-term model choice.

Where o4-mini fits in OpenAI's lineup

o4-mini is part of OpenAI's o-series reasoning family. Its role is not to provide the broadest possible set of media-generation capabilities or the maximum depth of reasoning. Instead, it targets workloads that need reasoning repeatedly, at relatively high volume, with controlled cost and latency.

That positioning matters when choosing between model types. A standard lightweight language model may be preferable for simple classification, rewriting, or straightforward extraction when deliberate reasoning is unnecessary. A larger reasoning model may be more appropriate for unusually difficult problems where answer quality is more important than response time and token cost. o4-mini is most attractive when the workload sits between those extremes.

OpenAI identifies GPT-5 Mini as the successor to o4-mini in its current API documentation. That does not mean every application will behave identically after migration. Teams should test prompts, tool calls, structured responses, latency, and output quality before replacing o4-mini.

Inputs, outputs, and supported capabilities

o4-mini accepts text and image input and returns text output. Images can therefore be included in tasks such as interpreting charts, answering questions about screenshots, reviewing diagrams, or extracting information from visual documents. The model does not natively accept audio or video, and it does not generate images, audio, or video.

The model supports reasoning tokens, which are internal tokens used while working through a problem. They contribute to usage and output limits even though they are not necessarily presented as ordinary answer text. For applications, the practical result is that a request can involve more computation than the visible response alone suggests.

Supported API features include:

  • Streaming responses, allowing an application to receive output progressively.
  • Function calling, so the model can request that an application run an external function or tool.
  • Structured outputs for responses that must follow a defined format.
  • Batch processing for workloads that can be submitted and processed outside an interactive request flow.
  • Reinforcement fine-tuning for the documented dated o4-mini snapshot, enabling customization for certain domain-specific behaviors.

In supported Responses API workflows, o4-mini can use hosted web search. Web search can provide retrieved information during a request, but it does not alter the model's underlying June 1, 2024 knowledge cutoff. Applications that need current information should therefore use retrieval or tools deliberately rather than relying on the model's built-in knowledge.

Context window and output limits

o4-mini has a 200,000-token context window. The context window is the total amount of information the request and response process can accommodate, including supplied text, images as represented by the API, conversation history, tool material, and generated content.

The maximum output is 100,000 tokens. This is a ceiling rather than a recommended response size; most normal answers, code reviews, and extraction tasks should use far fewer tokens. Long outputs can increase cost and may add unnecessary latency, so applications should set practical output limits that match the task.

For web-search workflows, OpenAI's documentation specifies a 128,000-token limit for web-search context. This is distinct from the model's general 200,000-token context window and is an important constraint for applications that combine large prompts with retrieved search material.

Pricing and cost trade-offs

Standard API pricing is $1.10 per 1 million input tokens and $4.40 per 1 million output tokens. Cached input tokens cost $0.275 per 1 million tokens. Cached-input pricing can matter for applications that repeatedly send an unchanged system prompt, reference material, or other reusable context.

Input and output prices are separate. A response that generates substantial reasoning or a long visible answer can therefore cost more on the output side, while a large document, conversation history, or tool result primarily affects input usage. Tool calls may have separate charges, so the model's token prices should not be treated as the complete cost of a tool-enabled workflow.

o4-mini's cost advantage is most useful at scale. For example, it can be a practical choice for automated code checks, mathematical classification, image-based extraction, or agent steps that run across many requests. Its value is less clear when a task is simple enough for a conventional low-cost model or so demanding that the application needs the strongest available reasoning quality.

Reasoning, coding, and tool use

Reasoning is the model's defining capability. It is designed for multi-step tasks rather than only predicting a short continuation. In practical terms, that makes it suitable for breaking down mathematical problems, following technical constraints, comparing alternatives, and checking intermediate logic.

For coding, o4-mini can generate code, explain implementation choices, review existing code, identify likely defects, and help with multi-step technical tasks. It is especially relevant when a coding request includes several constraints or requires interpreting an image such as a diagram, interface screenshot, or chart. As with any coding model, generated code should be tested rather than accepted solely because the explanation sounds plausible.

Function calling allows the model to work as one component in an application rather than as an isolated chatbot. The model can decide that an external operation is needed, provide arguments in the expected format, and use the returned result as part of its next response. This is useful for database lookups, business workflows, calculations, and retrieval systems. The application remains responsible for validating arguments, enforcing permissions, executing functions safely, and handling failures.

Structured outputs are useful when the result must be consumed by software. A developer can request a defined response structure for tasks such as classification, extraction, routing, or form processing. Structured output support should not automatically be confused with a separate legacy JSON-mode capability; the supplied documentation verifies structured outputs but does not independently verify that distinct capability for o4-mini.

Best use cases for o4-mini

o4-mini is a strong fit when the application needs reasoning repeatedly but must control cost or response time. Suitable workloads include:

  • Mathematical, scientific, and data-analysis assistance.
  • Code generation, debugging support, and automated code review.
  • Visual question answering involving screenshots, charts, or diagrams.
  • Structured extraction from text and images.
  • Technical support that requires following several diagnostic steps.
  • High-volume assistants and agents that call external tools.
  • Classification or routing tasks where the decision requires more than simple keyword matching.

The model is particularly useful when a request combines modalities or capabilities, such as asking it to inspect a chart, reason about the result, and return a structured summary. It can also serve as a lower-cost reasoning step inside a larger workflow that uses retrieval, business functions, or validation code.

Limitations and when to choose another model

o4-mini is not a general solution for every media or knowledge requirement. It does not accept audio or video as native inputs and does not produce image, audio, or video output. Applications centered on speech, video understanding, or media generation need a model with those modalities or a separate processing pipeline.

Its June 1, 2024 knowledge cutoff also limits its reliability for events, products, regulations, or technical information that changed afterward. Web search or another retrieval system can help, but retrieval must be implemented and evaluated as part of the application.

The model is also a poor long-term choice for a new integration if the team cannot complete a migration before the scheduled API shutdown on October 23, 2026. OpenAI has already marked it deprecated, and it retired o4-mini from ChatGPT on February 13, 2026 even though API access continued afterward.

Choose o4-mini when fast, cost-sensitive reasoning is more important than maximum depth and the task mainly involves text or images. Choose a simpler non-reasoning model when the task is routine and latency or price dominates. Consider a larger or newer reasoning model when the problem is unusually difficult, current model support is essential, or the application requires a longer future lifecycle. For existing o4-mini deployments, GPT-5 Mini is the named successor documented by OpenAI, but migration should be verified with representative workloads rather than assumed to be drop-in compatible.

Availability and migration planning

o4-mini remains available through OpenAI's API in the supplied documentation, including the Chat Completions and Responses APIs. However, its deprecated status and scheduled shutdown make lifecycle planning part of the technical evaluation.

Before migration, record representative prompts and expected outputs, including image inputs, structured responses, function-call arguments, long-context requests, and failure cases. Compare answer quality, token use, latency, tool behavior, and validation requirements with the replacement model. Applications should also avoid hard-coding assumptions about the unversioned alias and dated snapshot, because their lifecycle status and behavior may not remain identical.

In summary, o4-mini is best understood as a fast, economical reasoning model for text-and-image workloads, not as a permanent general-purpose endpoint. Its combination of reasoning, coding support, visual input, tool use, structured outputs, and high context capacity makes it useful today, while its deprecation and scheduled shutdown mean that new projects should include a clear replacement plan.


Answers to Frequently Asked Questions

Is o4-mini still available, and what is its successor?
o4-mini remains available through OpenAI's API in the supplied documentation, but it is deprecated and scheduled for API shutdown on October 23, 2026. OpenAI identifies GPT-5 Mini as its successor. Teams should test prompts, tool calls, structured outputs, latency, token usage, and response quality before migrating.
What are o4-mini's context window and output limits?
o4-mini has a 200,000-token context window and a maximum output of 100,000 tokens. For workflows using web search, the documented web-search context limit is 128,000 tokens.
How much does the o4-mini API cost?
Standard API pricing is $1.10 per 1 million input tokens, $4.40 per 1 million output tokens, and $0.275 per 1 million cached input tokens. Tool calls may incur additional charges, and reasoning tokens contribute to usage.
What inputs and outputs does o4-mini support?
o4-mini accepts text and image inputs and returns text. It can analyze screenshots, charts, diagrams, and visual documents, but it does not natively accept audio or video or generate image, audio, or video content.
What is OpenAI o4-mini used for?
o4-mini is a compact reasoning model designed for mathematics, coding, code review, visual question answering, structured extraction, data analysis, technical support, and high-volume API workflows that need more reasoning than a basic language model but lower cost and latency than larger reasoning models.


Sources 6
Provider

About OpenAI