GPT-4o Mini Search Preview

GPT-4o Mini Search Preview

by OpenAI · Retired; access shut down on July 23, 2026

GPT-4o Mini Search Preview was OpenAI’s lightweight Chat Completions model for automatic web-search queries. It offered a 128K-token context window, 16,384-token output limit, low token pricing, streaming, and structured outputs, but was text-only, lacked function calling and advanced search controls, and was shut down on July 23, 2026.

Text Reasoning Coding
GPT-4o Mini Search Preview was a specialized OpenAI model built to search the web before generating a text response. Released on March 11, 2025, it targeted applications that needed search-grounded answers at lower cost and potentially lower latency than larger search-oriented models. The model supported a 128,000-token context window, up to 16,384 output tokens, streaming, and structured outputs. It is now retired: OpenAI deprecated the model and shut down access on July 23, 2026.
Outputs

What GPT-4o Mini Search Preview can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Structured output Batch API
Model profile

Performance characteristics

5/10 Reasoning
4/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o Mini Search Preview
Model type Lightweight
Context window 128K tokens
Maximum output 16K tokens
Knowledge cutoff 2023-10-01
Release date 2025-03-11
Status Retired; access shut down on July 23, 2026
Deprecation date 2026-04-22
Shutdown date 2026-07-23
Knowledge cutoff notes

The official model page lists October 1, 2023 as the underlying knowledge cutoff. Web search could provide newer information during model use, but did not change the model's documented cutoff.

Model notes

Specialized search model for the Chat Completions API rather than a general-purpose multimodal model. The canonical alias and dated snapshot were deprecated; access to gpt-4o-mini-search-preview-2025-03-11 ended on July 23, 2026. OpenAI recommends migrating to the Responses API web_search tool or gpt-5-search-api. Web search incurred an additional per-call fee. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input $0.15 per 1 million input tokens, plus a separate fee per web-search tool call
Output $0.60 per 1 million output tokens
Model guide

GPT-4o Mini Search Preview: OpenAI’s Retired Low-Cost Web Search Model

GPT-4o Mini Search Preview was OpenAI’s lightweight, text-only model for automatic web-search workflows in the Chat Completions API. It combined a 128,000-token context window with low token pricing and fast operation, but lacked function calling, offered fewer search controls than the newer Responses API web_search tool, and was shut down on July 23, 2026.

What GPT-4o Mini Search Preview was

GPT-4o Mini Search Preview was OpenAI’s small, specialized model for web-search queries. It was designed for the legacy Chat Completions API, where the search model handled the search step before producing a response. This made it different from a conventional general-purpose language model: its main reason for existing was to generate answers grounded in current web information.

OpenAI introduced it on March 11, 2025, alongside GPT-4o Search Preview and computer-use-preview. The “Mini” designation reflected its position as the smaller, lower-cost option in OpenAI’s search-model offering. It was intended for applications where search-grounded text mattered more than multimodal input, advanced tool orchestration, or maximum reasoning capability.

The model is no longer available. OpenAI marked it as deprecated, and access to both the canonical model and the dated snapshot gpt-4o-mini-search-preview-2025-03-11 ended on July 23, 2026.

Specifications and pricing

SpecificationDocumented value
ProviderOpenAI
Model IDgpt-4o-mini-search-preview
Release dateMarch 11, 2025
StatusRetired; access shut down July 23, 2026
Context window128,000 tokens
Maximum output16,384 tokens
Knowledge cutoffOctober 1, 2023
Input modalityText
Output modalityText
Primary APIChat Completions

Token pricing was $0.15 per 1 million input tokens and $0.60 per 1 million output tokens. OpenAI also documented a separate fee for each web-search tool call. The token rates made the model suitable for high-volume text workloads, but the total cost of a request depended on both token usage and how often the application triggered web search.

The 128,000-token context window allowed an application to provide a substantial amount of conversation history or source material in one request. The maximum generated response was 16,384 tokens. These limits describe the model’s documented capacity; they do not mean that every practical request should use the full window or produce a very long answer.

How its search workflow worked

GPT-4o Mini Search Preview was trained to understand and execute web-search queries. In the legacy search-model workflow, an application sent a request to Chat Completions using the search model, and the model performed search before generating its response. Search was therefore closely associated with the model rather than exposed as the independently configured tool used by newer OpenAI APIs.

This approach was convenient for applications that simply needed a current, search-grounded answer. For example, a service could use the model to answer a question about a recent event, summarize information discovered online, or retrieve current web information before returning a text response to a user.

However, the legacy workflow had fewer controls than the newer Responses API web_search tool. The supplied documentation identifies limitations involving domain filtering, complete source lists, live-access control, and returned-token budget control. Applications requiring detailed control over how search is performed or how sources are exposed would therefore have been better suited to the newer tool-based approach.

Modalities and API features

GPT-4o Mini Search Preview accepted text input and generated text output. It was not documented as an image, audio, or video model. Web search was its defining specialized capability, but that did not make it a general multimodal system.

The documented feature set included streaming and structured outputs. Streaming allows an application to receive generated text incrementally instead of waiting for the complete response. Structured outputs can help an application request data that follows a defined structure, although the supplied research does not establish a separate JSON-mode capability for this model.

Function calling was not supported. This is an important distinction: the model could perform its specialized web-search workflow, but it was not a general-purpose function-calling model that could reliably select and invoke arbitrary application-defined functions. Systems that needed database actions, business-logic calls, or multi-step tool orchestration needed another option.

Strengths and trade-offs

The model’s main strength was specialization. It combined automatic web search with relatively low input and output token prices, making it a reasonable fit for applications that needed current information without using a larger, more expensive model for every request. Its 128,000-token context window also gave it room to process long prompts or substantial retrieved content.

Its documented feature profile suggests a speed-and-cost trade-off rather than a focus on frontier reasoning. Editorial comparisons in the supplied research rate its speed highly and its cost favorably, while assigning more modest scores to reasoning and coding. Those scores are comparative editorial estimates, not OpenAI benchmarks or vendor-published measurements. They should be read as guidance about positioning, not as guaranteed performance levels.

The same specialization created limitations. The model was text-only, did not support function calling, and did not provide the richer search configuration available through the Responses API web-search tool. Its underlying knowledge cutoff was October 1, 2023. Web search could supply newer information during a request, but it did not update the model’s underlying training knowledge or remove the need to assess the quality and relevance of retrieved information.

When to choose this model

While it was available, GPT-4o Mini Search Preview was most appropriate for legacy Chat Completions applications with these characteristics:

  • The application needed current, web-grounded text responses rather than offline generation.
  • Search should occur automatically as part of the model request.
  • Low token cost and high throughput mattered more than advanced reasoning or broad modality support.
  • The workflow did not require application-defined function calling.
  • The application could work with the search controls and source behavior of the legacy search-model path.

Typical examples included a high-volume question-answering service, a simple current-information assistant, or a text summarization workflow that needed the model to search before responding.

There is no longer a valid production case for selecting this model because it was shut down. Existing designs should not treat the old model ID as a dependable fallback or assume that a replacement will behave as an alias.

When another option is more appropriate

For new integrations, OpenAI recommends the Responses API web_search tool or gpt-5-search-api, depending on the required integration style. The Responses API option is more appropriate when an application needs the newer tool-based search workflow and its richer controls. GPT-5 Search API is the documented direction when continued use of a Chat Completions search-model path is required.

A general-purpose model is a better fit when web search is not central to the task or when the application needs capabilities that this model did not provide, such as function calling or multimodal input. A larger or more capable model may also be preferable for complex reasoning, although the supplied research does not provide a direct benchmark comparison with GPT-4o Mini Search Preview.

Applications that need images, audio, or video should choose a model explicitly documented for those modalities. GPT-4o Mini Search Preview accepted and returned text only, so adding a search requirement does not change that limitation.

Retirement and migration

OpenAI announced the model’s deprecation on April 22, 2026, with shutdown scheduled for July 23, 2026. The supplied current-status information records that the shutdown occurred on that date. As a result, requests using gpt-4o-mini-search-preview should be migrated rather than debugged as though the model were temporarily unavailable.

Migration requires a design decision, not just a string replacement. The Responses API web_search tool represents search separately from the model and offers a different control surface. The GPT-5 Search API is a separate search-model option. Developers should review request formats, source handling, tool behavior, pricing, and output expectations before moving production traffic.

In historical terms, GPT-4o Mini Search Preview was useful because it made low-cost web-grounded text generation available through a specialized Chat Completions model. In practical terms today, its most important fact is its retirement: it can explain the legacy search-model pattern, but it should not be selected for a new application.


Answers to Frequently Asked Questions

What should developers use instead of GPT-4o Mini Search Preview?
For new integrations, developers should consider the Responses API web_search tool or gpt-5-search-api, depending on whether they need a tool-based workflow or a Chat Completions search-model path. Migration requires reviewing request formats, source handling, tool behavior, pricing, and output expectations.
What were the main limitations of GPT-4o Mini Search Preview?
The model accepted and generated text only, did not support function calling, and offered fewer search controls than the newer Responses API web_search tool. Its underlying knowledge cutoff was October 1, 2023, although web search could retrieve newer information during a request.
How much did GPT-4o Mini Search Preview cost?
The model cost $0.15 per 1 million input tokens and $0.60 per 1 million output tokens. OpenAI also charged a separate fee for each web-search tool call, so total request costs depended on both token usage and search frequency.
What was GPT-4o Mini Search Preview?
GPT-4o Mini Search Preview was OpenAI’s low-cost, text-only search model for the legacy Chat Completions API. It performed web searches before generating responses grounded in current online information.
Is GPT-4o Mini Search Preview still available?
No. OpenAI deprecated GPT-4o Mini Search Preview, including the snapshot gpt-4o-mini-search-preview-2025-03-11, and shut down access on July 23, 2026.


Sources 5
Provider

About OpenAI