GPT-5

GPT-5 Mini

by OpenAI · Current API alias; dated snapshot gpt-5-mini-2025-08-07 is deprecated

OpenAI's GPT-5 Mini is a cost-efficient reasoning model for high-volume API workloads. It supports text and image input, text output, a 400,000-token context window, 128,000-token responses, function calling, structured outputs, streaming, prompt caching, and Batch API processing.

Text Reasoning Coding
GPT-5 Mini is a faster and more cost-efficient member of OpenAI's GPT-5 family. It is designed for applications that need capable reasoning and coding at lower cost and latency than a larger frontier model, while retaining a large context window and support for image input.
Outputs

What GPT-5 Mini can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family GPT-5
Model type Lightweight
Context window 400K tokens
Maximum output 128K tokens
Knowledge cutoff 2024-05-31
Release date 2025-08-07
Status Current API alias; dated snapshot gpt-5-mini-2025-08-07 is deprecated
Knowledge cutoff notes

The official GPT-5 Mini model page explicitly states a May 31, 2024 knowledge cutoff. Web search, retrieval, or other external context can provide newer information during use but do not change the underlying cutoff.

Model notes

GPT-5 Mini is the faster, lower-cost GPT-5 variant for well-defined tasks and precise prompts. The current canonical alias is gpt-5-mini. The dated snapshot gpt-5-mini-2025-08-07 is listed as deprecated. OpenAI documents text and image input, text output, reasoning token support, streaming, function calling, structured outputs, prompt caching, and Batch API access. The model page lists a 400,000-token context window, 128,000-token maximum output, and a May 31, 2024 knowledge cutoff. OpenAI recommends GPT-5.6 Terra for most new low-latency, high-volume workloads, but GPT-5 Mini remains listed as an accessible model alias.

Cost

Model pricing

Input US$0.25 per 1 million input tokens; cached input US$0.025 per 1 million tokens
Output US$2.00 per 1 million output tokens
Model guide

GPT-5 Mini: Features, Pricing, Context Window, and API Availability

GPT-5 Mini is OpenAI's lower-cost reasoning model for well-defined, high-volume workloads. It accepts text and images, produces text, supports a 400,000-token context window, up to 128,000 output tokens, structured outputs, streaming, function calling, prompt caching, and Batch API processing.

What is GPT-5 Mini?

GPT-5 Mini is OpenAI's compact GPT-5 reasoning model for applications that need to process many requests without the cost of the full GPT-5 model. OpenAI positions it for well-defined tasks, precise prompts, coding, analysis, instruction following, and tool-enabled workflows.

The model is available through OpenAI's API under the canonical model ID gpt-5-mini. It was released to the API on August 7, 2025. OpenAI also lists the dated snapshot gpt-5-mini-2025-08-07, but that snapshot is deprecated; new integrations should generally use the current undated alias when it meets the application's requirements.

In practical terms, GPT-5 Mini occupies the part of the GPT-5 lineup where cost, response speed, and throughput matter more than maximum available reasoning quality. It is not a consumer ChatGPT subscription or a separate ChatGPT plan. The prices below refer to API usage.

Capabilities and supported modalities

GPT-5 Mini accepts text and image input and returns text. Image input allows an application to provide visual material for tasks such as document understanding or image-based analysis, but the model does not generate images. It also does not natively accept audio or video and does not produce audio or video.

  • Input: Text and images
  • Output: Text only
  • Context window: 400,000 tokens
  • Maximum output: 128,000 tokens
  • Reasoning: Supports reasoning tokens for multi-step problem solving
  • Tools: Function calling and supported tool workflows
  • Generation controls: Streaming and structured outputs
  • Efficiency features: Prompt caching and Batch API processing

A token is a small unit of text used for billing and model context. The 400,000-token context window is the amount of combined prompt, supplied documents, conversation history, and other input the model can consider in one request. The 128,000-token maximum output is a separate ceiling for the response. Actual usable limits can also depend on the request configuration and the API surface.

Reasoning and coding performance

GPT-5 Mini is intended for tasks where the model must follow several instructions, inspect information, or reach an answer through multiple steps rather than simply continue text. Reasoning tokens are internal processing tokens used during this kind of work; they are distinct from the visible answer returned to the application.

That design makes the model suitable for structured extraction, classification that requires judgment, document analysis, coding assistance, and agents that need to decide which tool or function to call. It can also generate code and help analyze or transform existing code, although the supplied research does not establish that it is the best choice for every software-engineering task.

The model's reasoning and coding ability should be understood as a trade-off rather than an absolute ranking. OpenAI describes GPT-5 Mini as a lower-cost, lower-latency option for well-defined workloads. A larger or newer model may be more appropriate when a task is unusually ambiguous, requires the highest level of reasoning quality, or benefits from frontier performance more than from throughput and price efficiency.

Context window and knowledge cutoff

The 400,000-token context window is one of GPT-5 Mini's most useful specifications. It can provide room for long documents, sizeable code repositories, retrieval results, or multi-step prompts without dividing all input into many smaller requests. This does not mean the model automatically knows every document or understands every long prompt perfectly; the relevant material still needs to be supplied and the application should validate important results.

OpenAI documents a knowledge cutoff of May 31, 2024. Consequently, GPT-5 Mini should not be treated as a current-events database. For information that may have changed since that date, an application should provide current material through retrieval, web search, a connected data source, or another external context. The model's support for tool-enabled workflows can help applications do this, but external tools do not change the underlying training cutoff.

GPT-5 Mini API pricing

OpenAI's listed pricing is:

Token categoryPrice
Input tokensUS$0.25 per 1 million tokens
Cached input tokensUS$0.025 per 1 million tokens
Output tokensUS$2.00 per 1 million tokens

Input tokens are the material sent to the model, while output tokens are generated in the response. Cached input pricing applies when eligible repeated prompt content is handled through prompt caching; it is not a general discount on every request. Output tokens cost more than ordinary input tokens, so applications can reduce expenditure by keeping instructions concise, reusing stable prompt prefixes where appropriate, and requesting only the response detail they need.

GPT-5 Mini is also intended for high-volume processing through the Batch API. Batch processing is useful when results do not need to be returned immediately, but the supplied research does not provide a separate Batch price, so the standard listed prices should not be replaced with an invented estimate.

Tools, structured outputs, and API support

GPT-5 Mini is available through both the Responses API and the Chat Completions API. Function calling lets an application describe functions that the model can request, such as searching a database, checking an order, or starting a workflow. The application—not the model—executes the function and returns the result, allowing the model to incorporate that information into a final response.

Structured outputs are useful when the response must conform to a defined schema instead of being free-form prose. For example, an extraction workflow can request fields for a document's invoice number, date, supplier, and total. The model supports streaming as well, allowing an application to display or process generated text progressively rather than waiting for the entire response.

OpenAI documents web search and other built-in tool workflows when they are configured through supported OpenAI APIs. Tool availability depends on the API configuration and application design; the model should not be assumed to have unrestricted browsing or access to a user's private systems by default.

Main strengths and limitations

GPT-5 Mini's strongest combination is its relatively low API price, high throughput orientation, large context window, and support for reasoning and tools. That combination is valuable when an application processes many documents, creates structured records, routes support requests, assists with code, or runs repeated agent steps.

  • Large context: 400,000 tokens can accommodate long inputs and extensive supporting material.
  • Low listed cost: Input and output pricing is substantially below what is normally expected from a top-tier frontier model, making repeated workloads easier to scale.
  • Good workflow coverage: Function calling, structured outputs, streaming, caching, and Batch API support cover common production patterns.
  • Vision input: Images can be supplied alongside text for supported analysis tasks.
  • Reasoning and coding: The model is designed for multi-step analysis, coding help, and precise instruction following.

There are also clear boundaries. GPT-5 Mini produces text only, so it is not the correct model for native audio or video processing or for image generation. It has a May 31, 2024 knowledge cutoff and therefore needs external information for current facts. It is positioned below larger or newer models when maximum reasoning quality is the priority. Finally, model output can still be incorrect or overconfident, so applications should validate critical calculations, extracted fields, code, and decisions.

When to choose GPT-5 Mini

Choose GPT-5 Mini when the task is sufficiently defined to benefit from a repeatable prompt and the application needs a balance of reasoning, speed, context capacity, and cost. Suitable examples include:

  • High-volume customer-support responses with function calls to account or order systems
  • Document classification, extraction, and conversion into structured records
  • Long-document question answering when the source material is supplied in the prompt or through retrieval
  • Code explanation, code transformation, test generation, and routine development assistance
  • Retrieval-augmented generation systems that need to combine multiple source passages
  • Agent workflows that repeatedly call tools and return structured results
  • Batch processing where immediate responses are not required

It is less suitable when the application needs audio or video input, generated images, or a text-to-speech or speech-recognition model. It is also a weaker fit when a difficult, ambiguous problem justifies paying for a larger frontier model. In those cases, compare the task against a full-size GPT-5 option such as GPT-5, rather than assuming the Mini variant will deliver identical reasoning quality at a lower price.

For a newer low-latency, high-volume workload, OpenAI currently recommends GPT-5.6 Terra in the supplied model guidance. That recommendation does not make GPT-5 Mini unavailable or interchangeable with GPT-5.6 Terra: each model has its own behavior, pricing, limits, and lifecycle. Teams should test representative prompts and compare total request cost, latency, output quality, tool reliability, and context requirements before migrating.

Current API status

The current gpt-5-mini alias is listed in OpenAI's API model catalog. The dated snapshot gpt-5-mini-2025-08-07 is deprecated, so it should not be selected for new integrations when the current alias is appropriate. Applications that require stable behavior should still monitor OpenAI's lifecycle documentation and test model changes, because an alias can represent the provider's current version rather than permanently pinning a dated snapshot.

Overall, GPT-5 Mini is best understood as a cost-conscious reasoning model for production workloads with clear requirements. Its combination of text and image input, large context, structured responses, function calling, and low listed token prices makes it practical for scalable automation. Its text-only output, dated knowledge cutoff, and lower positioning within the GPT-5 family are equally important when deciding whether it fits a particular application.


Answers to Frequently Asked Questions

Is GPT-5 Mini available through the OpenAI API?
Yes. GPT-5 Mini is available through both the Responses API and Chat Completions API under the model ID "gpt-5-mini". The dated snapshot "gpt-5-mini-2025-08-07" is deprecated, so new integrations should generally use the current undated alias.
What modalities and tools does GPT-5 Mini support?
GPT-5 Mini accepts text and image input and produces text output. It supports reasoning tokens, function calling, structured outputs, streaming, prompt caching, and Batch API processing. It does not natively accept or generate audio or video and cannot generate images.
How much does GPT-5 Mini API access cost?
GPT-5 Mini costs US$0.25 per 1 million input tokens, US$0.025 per 1 million cached input tokens, and US$2.00 per 1 million output tokens. Cached pricing applies only to eligible repeated prompt content.
What are GPT-5 Mini's context window and output limits?
GPT-5 Mini has a 400,000-token context window and a maximum output limit of 128,000 tokens. The context window includes the prompt, documents, conversation history, and other supplied input, while the output limit applies only to the generated response.
What is GPT-5 Mini?
GPT-5 Mini is OpenAI's compact GPT-5 reasoning model for high-volume applications that need lower cost and latency than the full GPT-5 model. It is designed for well-defined tasks, coding, analysis, instruction following, and tool-enabled workflows through the API.


Sources 4
Provider

About OpenAI