GPT-4o

GPT-4o Search Preview

by OpenAI · Retired; shut down on 2026-07-23

GPT-4o Search Preview was OpenAI’s dedicated web-search model for Chat Completions. It offered a 128,000-token context window, 16,384-token maximum output, streaming, and structured outputs, with historical pricing of $2.50 per million input tokens and $10 per million output tokens. It was text-only, lacked function calling and fine-tuning, and was shut down on July 23, 2026.

Text Reasoning Coding
GPT-4o Search Preview was built for applications that needed current information from the web rather than relying only on a model’s fixed training data. The OpenAI model searched before generating an answer through Chat Completions, with a large context window and support for streaming and structured outputs. Its limitations were equally important: it accepted and returned text only, lacked general function calling and fine-tuning, and is no longer available because OpenAI retired it on July 23, 2026.
Outputs

What GPT-4o Search Preview can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Structured output
Model profile

Performance characteristics

5/10 Reasoning
5/10 Coding
6/10 Speed
3/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o
Model type Other
Context window 128K tokens
Maximum output 16K tokens
Knowledge cutoff 2023-10-01
Release date 2025-03-11
Status Retired; shut down on 2026-07-23
Deprecation date 2026-07-23
Shutdown date 2026-07-23
Knowledge cutoff notes

OpenAI lists October 1, 2023 as the model's underlying knowledge cutoff. Web search supplied current external information during model use but did not change the underlying cutoff.

Model notes

GPT-4o Search Preview was a specialized search model rather than a general-purpose GPT-4o endpoint. Its canonical alias was gpt-4o-search-preview and its dated snapshot was gpt-4o-search-preview-2025-03-11. OpenAI listed a 128,000-token context window, 16,384 maximum output tokens, text-only input and output, streaming, and structured outputs. Function calling and fine-tuning were not supported. The model was deprecated and shut down on July 23, 2026. OpenAI recommends migrating to the Responses API web_search tool or GPT-5 Search API. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input $2.50 per 1 million input tokens; historical web-search tool fees applied separately per search call
Output $10.00 per 1 million output tokens
Model guide

GPT-4o Search Preview: OpenAI’s Dedicated Web-Search Model

GPT-4o Search Preview was OpenAI’s specialized web-search model for the Chat Completions API. It combined automatic web searching with text generation, supported a 128,000-token context window, streaming, and structured outputs, and charged $2.50 per million input tokens plus $10 per million output tokens. It was text-only, did not support ordinary function calling or fine-tuning, and was deprecated and shut down on July 23, 2026.

What was GPT-4o Search Preview?

GPT-4o Search Preview was a specialized OpenAI model designed to search the web before producing an answer. It was intended for applications using the Chat Completions API that needed timely external information, such as recent news, changing product details, or answers that depended on current web content.

The model’s web-search behavior was separate from its underlying training knowledge. OpenAI listed October 1, 2023 as the knowledge cutoff for the model itself, while web search provided information from external sources during a request. This distinction matters: the model was not continually retrained as the web changed, but it could retrieve current information while answering supported queries.

GPT-4o Search Preview launched on March 11, 2025. Its canonical model alias was gpt-4o-search-preview, and OpenAI also published the dated snapshot gpt-4o-search-preview-2025-03-11.

Where it fit in OpenAI’s catalog

GPT-4o Search Preview was not simply a general-purpose GPT-4o endpoint with an optional search switch. It was a dedicated search-oriented model optimized around the task of finding web information and using it in a generated response. That specialization made it more relevant to current-information applications than to broad multimodal or tool-orchestration workloads.

Its API design also reflected an earlier stage of OpenAI’s search tooling. OpenAI has recommended that users migrate legacy integrations to the Responses API’s web_search tool, or use the GPT-5 Search API when Chat Completions-style search-model compatibility is required. These alternatives are migration destinations, not capabilities that should be attributed to GPT-4o Search Preview itself.

Technical specifications and limits

The following are the documented specifications supplied for GPT-4o Search Preview. The token figures describe the model’s context and output capacity: a token is a small unit of text used by language models for processing and billing.

SpecificationDocumented value
ProviderOpenAI
Release dateMarch 11, 2025
Context window128,000 tokens
Maximum output16,384 tokens
Input and outputText only
StreamingSupported
Structured outputsSupported
Function callingNot supported
Fine-tuningNot supported
StatusDeprecated and shut down July 23, 2026

The 128,000-token context window allowed an application to provide a substantial amount of text alongside a request. The maximum generated response was 16,384 tokens. These limits did not make the model multimodal: image, audio, and video inputs were not supported, and the model did not directly generate images, audio, or video.

Web search and API behavior

Search was the model’s defining capability. Rather than asking a general language model to answer only from its static knowledge, an application could use GPT-4o Search Preview for requests where online information was important. The model searched the web and then generated a text response based on the resulting information.

Search requests incurred an additional per-call fee on top of token charges. The supplied research does not specify the amount of that historical search fee, so it should not be estimated from the token prices. The model supported streaming, which allowed an application to receive a response incrementally instead of waiting for the complete answer.

GPT-4o Search Preview did not support ordinary function calling. Function calling lets a model select and populate application-defined functions, such as a database lookup or an order-status request. The absence of that feature limited the model’s usefulness in systems that needed coordinated workflows beyond web search and text generation. Its built-in search specialization should therefore not be confused with general-purpose tool orchestration.

Pricing and cost trade-offs

The documented historical token prices were:

  • Input: $2.50 per 1 million tokens
  • Output: $10.00 per 1 million tokens
  • Web search: an additional fee per search call

Input tokens are the text sent to the model, while output tokens are the text it generates. Output was priced higher than input, so applications that requested unnecessarily long answers could increase their token cost. Search added another cost layer, making the model’s economics different from those of a text model that answers without retrieval.

These are historical prices for a retired model, not a current purchasing option. Because GPT-4o Search Preview has been shut down, new production systems should evaluate the pricing of OpenAI’s recommended replacement rather than budget around this endpoint.

Main strengths and limitations

Strengths

  • Specialized current-information workflow: web search was central to the model’s design rather than an improvised add-on.
  • Large context window: the 128,000-token capacity supported substantial prompts and retrieved material.
  • Longer response ceiling: up to 16,384 output tokens were documented.
  • API-friendly output options: streaming and structured outputs were supported.
  • Chat Completions compatibility: it was designed for applications already using that API pattern.

Limitations

  • Retired status: the endpoint was shut down on July 23, 2026 and cannot be selected for new production use.
  • Text-only operation: it did not accept images, audio, or video and did not produce non-text media.
  • No ordinary function calling: it was not a general tool-orchestration model.
  • No fine-tuning: organizations could not fine-tune the model for a specialized behavior.
  • Additional search charges: search calls cost extra beyond input and output token usage.
  • Historical knowledge cutoff: the underlying model knowledge cutoff was October 1, 2023; current information depended on successful web retrieval.

Reasoning, coding, speed, and cost

OpenAI’s supplied specifications do not identify GPT-4o Search Preview as a dedicated reasoning model, and they do not provide standardized public benchmark results for reasoning or coding. Editorial comparative scores in the research rate reasoning and coding at 5 out of 10 and speed at 6 out of 10, but these are subjective evaluations rather than OpenAI-published benchmarks.

In practical terms, the model’s main advantage was search specialization, not advanced reasoning, software engineering, or multimodal understanding. Its speed and cost profile also depended on the search step, the amount of retrieved and supplied text, the generated response length, and the additional per-search fee. A smaller or non-search model could be more appropriate when the information was already available in an application’s data or when minimizing retrieval overhead mattered. A model with broader tool support would be preferable for workflows involving business systems, custom functions, or multiple external actions.

When to choose this model

When it was available, GPT-4o Search Preview was a reasonable fit for a narrowly defined class of applications:

  • Chat Completions integrations that needed answers based on current web information.
  • Text-based research assistants and question-answering systems.
  • Applications that benefited from streaming responses or structured output formats.
  • Use cases where a 128,000-token context window was useful for combining a substantial prompt with retrieved material.

It was not a suitable choice for a new deployment today because the model has been retired. It was also a poor fit for image, audio, or video workloads; systems requiring function calls to application services; fine-tuned model deployments; or general-purpose multimodal assistants.

Retirement and migration

OpenAI deprecated GPT-4o Search Preview and shut it down on July 23, 2026. Existing integrations using gpt-4o-search-preview therefore need to be migrated rather than merely reconfigured.

OpenAI’s documented migration direction is the Responses API with its web_search tool. When an application specifically needs a search model compatible with the Chat Completions approach, OpenAI recommends the GPT-5 Search API. The right replacement depends on whether the priority is modern tool use through Responses or continued compatibility with a search-focused Chat Completions model.

Bottom line

GPT-4o Search Preview was a transitional OpenAI model that brought dedicated web search to Chat Completions. Its 128,000-token context window, streaming, structured outputs, and text-search specialization made it useful for current-information applications. However, it was limited to text, lacked ordinary function calling and fine-tuning, carried additional search costs, and is now shut down. It is best understood as a historical endpoint whose supported use cases have moved to OpenAI’s newer search tools and APIs.


Answers to Frequently Asked Questions

Is GPT-4o Search Preview still available?
No. OpenAI deprecated GPT-4o Search Preview and shut it down on July 23, 2026. Existing integrations should migrate to the Responses API with the web_search tool or, when Chat Completions compatibility is required, the GPT-5 Search API.
How much did GPT-4o Search Preview cost?
Its historical token pricing was $2.50 per 1 million input tokens and $10.00 per 1 million output tokens. Web searches incurred an additional per-call fee. These prices are historical because the model has been retired.
What were the main technical specifications of GPT-4o Search Preview?
The model had a 128,000-token context window, a maximum output of 16,384 tokens, text-only input and output, streaming support, and structured outputs. It did not support ordinary function calling, fine-tuning, or image, audio, and video inputs.
What was GPT-4o Search Preview?
GPT-4o Search Preview was a specialized OpenAI model that searched the web before generating answers. It was designed for Chat Completions API applications requiring current information, such as recent news, changing product details, and web-based research.


Sources 5
Provider

About OpenAI