What GPT-4o Mini Search Preview was
GPT-4o Mini Search Preview was OpenAI’s small, specialized model for web-search queries. It was designed for the legacy Chat Completions API, where the search model handled the search step before producing a response. This made it different from a conventional general-purpose language model: its main reason for existing was to generate answers grounded in current web information.
OpenAI introduced it on March 11, 2025, alongside GPT-4o Search Preview and computer-use-preview. The “Mini” designation reflected its position as the smaller, lower-cost option in OpenAI’s search-model offering. It was intended for applications where search-grounded text mattered more than multimodal input, advanced tool orchestration, or maximum reasoning capability.
The model is no longer available. OpenAI marked it as deprecated, and access to both the canonical model and the dated snapshot gpt-4o-mini-search-preview-2025-03-11 ended on July 23, 2026.
Specifications and pricing
| Specification | Documented value |
|---|---|
| Provider | OpenAI |
| Model ID | gpt-4o-mini-search-preview |
| Release date | March 11, 2025 |
| Status | Retired; access shut down July 23, 2026 |
| Context window | 128,000 tokens |
| Maximum output | 16,384 tokens |
| Knowledge cutoff | October 1, 2023 |
| Input modality | Text |
| Output modality | Text |
| Primary API | Chat Completions |
Token pricing was $0.15 per 1 million input tokens and $0.60 per 1 million output tokens. OpenAI also documented a separate fee for each web-search tool call. The token rates made the model suitable for high-volume text workloads, but the total cost of a request depended on both token usage and how often the application triggered web search.
The 128,000-token context window allowed an application to provide a substantial amount of conversation history or source material in one request. The maximum generated response was 16,384 tokens. These limits describe the model’s documented capacity; they do not mean that every practical request should use the full window or produce a very long answer.
How its search workflow worked
GPT-4o Mini Search Preview was trained to understand and execute web-search queries. In the legacy search-model workflow, an application sent a request to Chat Completions using the search model, and the model performed search before generating its response. Search was therefore closely associated with the model rather than exposed as the independently configured tool used by newer OpenAI APIs.
This approach was convenient for applications that simply needed a current, search-grounded answer. For example, a service could use the model to answer a question about a recent event, summarize information discovered online, or retrieve current web information before returning a text response to a user.
However, the legacy workflow had fewer controls than the newer Responses API web_search tool. The supplied documentation identifies limitations involving domain filtering, complete source lists, live-access control, and returned-token budget control. Applications requiring detailed control over how search is performed or how sources are exposed would therefore have been better suited to the newer tool-based approach.
Modalities and API features
GPT-4o Mini Search Preview accepted text input and generated text output. It was not documented as an image, audio, or video model. Web search was its defining specialized capability, but that did not make it a general multimodal system.
The documented feature set included streaming and structured outputs. Streaming allows an application to receive generated text incrementally instead of waiting for the complete response. Structured outputs can help an application request data that follows a defined structure, although the supplied research does not establish a separate JSON-mode capability for this model.
Function calling was not supported. This is an important distinction: the model could perform its specialized web-search workflow, but it was not a general-purpose function-calling model that could reliably select and invoke arbitrary application-defined functions. Systems that needed database actions, business-logic calls, or multi-step tool orchestration needed another option.
Strengths and trade-offs
The model’s main strength was specialization. It combined automatic web search with relatively low input and output token prices, making it a reasonable fit for applications that needed current information without using a larger, more expensive model for every request. Its 128,000-token context window also gave it room to process long prompts or substantial retrieved content.
Its documented feature profile suggests a speed-and-cost trade-off rather than a focus on frontier reasoning. Editorial comparisons in the supplied research rate its speed highly and its cost favorably, while assigning more modest scores to reasoning and coding. Those scores are comparative editorial estimates, not OpenAI benchmarks or vendor-published measurements. They should be read as guidance about positioning, not as guaranteed performance levels.
The same specialization created limitations. The model was text-only, did not support function calling, and did not provide the richer search configuration available through the Responses API web-search tool. Its underlying knowledge cutoff was October 1, 2023. Web search could supply newer information during a request, but it did not update the model’s underlying training knowledge or remove the need to assess the quality and relevance of retrieved information.
When to choose this model
While it was available, GPT-4o Mini Search Preview was most appropriate for legacy Chat Completions applications with these characteristics:
- The application needed current, web-grounded text responses rather than offline generation.
- Search should occur automatically as part of the model request.
- Low token cost and high throughput mattered more than advanced reasoning or broad modality support.
- The workflow did not require application-defined function calling.
- The application could work with the search controls and source behavior of the legacy search-model path.
Typical examples included a high-volume question-answering service, a simple current-information assistant, or a text summarization workflow that needed the model to search before responding.
There is no longer a valid production case for selecting this model because it was shut down. Existing designs should not treat the old model ID as a dependable fallback or assume that a replacement will behave as an alias.
When another option is more appropriate
For new integrations, OpenAI recommends the Responses API web_search tool or gpt-5-search-api, depending on the required integration style. The Responses API option is more appropriate when an application needs the newer tool-based search workflow and its richer controls. GPT-5 Search API is the documented direction when continued use of a Chat Completions search-model path is required.
A general-purpose model is a better fit when web search is not central to the task or when the application needs capabilities that this model did not provide, such as function calling or multimodal input. A larger or more capable model may also be preferable for complex reasoning, although the supplied research does not provide a direct benchmark comparison with GPT-4o Mini Search Preview.
Applications that need images, audio, or video should choose a model explicitly documented for those modalities. GPT-4o Mini Search Preview accepted and returned text only, so adding a search requirement does not change that limitation.
Retirement and migration
OpenAI announced the model’s deprecation on April 22, 2026, with shutdown scheduled for July 23, 2026. The supplied current-status information records that the shutdown occurred on that date. As a result, requests using gpt-4o-mini-search-preview should be migrated rather than debugged as though the model were temporarily unavailable.
Migration requires a design decision, not just a string replacement. The Responses API web_search tool represents search separately from the model and offers a different control surface. The GPT-5 Search API is a separate search-model option. Developers should review request formats, source handling, tool behavior, pricing, and output expectations before moving production traffic.
In historical terms, GPT-4o Mini Search Preview was useful because it made low-cost web-grounded text generation available through a specialized Chat Completions model. In practical terms today, its most important fact is its retirement: it can explain the legacy search-model pattern, but it should not be selected for a new application.

