What is o3-deep-research?
o3-deep-research is a specialized OpenAI reasoning model built for complex, multi-step investigations. Instead of focusing primarily on short conversational responses, it is intended to gather information, interpret documents, compare evidence, and synthesize findings into a research-oriented answer.
The model is accessed through OpenAI’s Responses API and can be combined with supported tools including web search, file search, remote Model Context Protocol (MCP) servers, and code interpreter workflows. MCP connections can provide access to external services or data sources, while file search can retrieve relevant material from connected vector stores. Code interpreter is useful when the research task includes calculations or computational analysis.
OpenAI released the model in June 2025 and currently marks the canonical o3-deep-research model as deprecated. That status is important for anyone considering a new production integration: the model may still explain an existing system or be useful for compatibility testing, but it is not the safest choice when long-term availability is a primary requirement.
Where it fits in OpenAI’s catalog
o3-deep-research occupies a specialized position within OpenAI’s model catalog. It is not presented as a general-purpose chat model, an image generator, or a conventional function-calling agent. Its role is to support research workflows in which the model must reason through multiple stages and work with information retrieved during the request.
This positioning creates a clear trade-off. A research-focused model can be more appropriate than a fast conversational model when the task depends on source coverage, lengthy documents, and evidence synthesis. The same design can be inefficient for simple questions, interactive chat, or applications where response time and predictable cost matter more than investigation depth.
The model should also be distinguished from a standalone web-search product. Its research behavior depends on the tools and data sources supplied through the Responses API. A request must include at least one supported data source, such as web search, a remote MCP server, or file search with vector stores. Tool availability therefore forms part of the practical implementation, rather than being an optional detail unrelated to the model.
Core capabilities and modalities
o3-deep-research accepts both text and image input and produces text output. Image input allows visual material to be considered alongside written instructions or documents, but the model does not natively generate images, audio, or video. Its output is written research, analysis, and explanation.
The model’s main capabilities include:
- Multi-step reasoning: It is intended for investigations that require planning, source interpretation, comparison, and synthesis.
- Web research: Supported web-search workflows can provide current information beyond the model’s underlying knowledge cutoff.
- File-based research: File search can retrieve relevant information from connected vector stores.
- External connections: Remote MCP servers can connect the research workflow to supported external data or services.
- Computational analysis: Code interpreter can be added when calculations or other code-based processing are needed.
- Long-form generation: The model supports a maximum output of 100,000 tokens, subject to the practical requirements and cost of the request.
These are tool-assisted research capabilities, not evidence that the model independently knows every current fact. Its official knowledge cutoff is June 1, 2024. Web search and connected sources can supply newer information during a request, but they do not change the model’s underlying cutoff.
Technical specifications
| Specification | Verified value |
|---|---|
| Provider | OpenAI |
| Model family | o3 |
| Model type | Reasoning |
| Context window | 200,000 tokens |
| Maximum output | 100,000 tokens |
| Knowledge cutoff | June 1, 2024 |
| Input | Text and image |
| Output | Text |
| Streaming | Supported |
| Function calling | Not supported |
| Structured outputs | Not supported |
| Fine-tuning | Not supported |
| API | Responses API |
A token is a unit of text used for processing and billing; it may represent a whole word, part of a word, punctuation, or other content. A 200,000-token context window provides room for substantial instructions, retrieved sources, and conversation history in one request. It does not mean every request should contain that much material: larger inputs can increase cost and may make source selection and organization more important.
Reasoning, coding, and tool support
OpenAI positions o3-deep-research for reasoning-heavy tasks rather than rapid response generation. In practical terms, this means it is a candidate for work such as comparing scientific literature, examining policy documents, investigating a market, or reviewing a large internal information set. The model’s value comes from coordinating research steps and turning collected material into a coherent result.
Code interpreter support extends the model beyond purely textual analysis. For example, a research workflow may use code to calculate statistics, transform a dataset, or perform other computational work needed to interpret evidence. The supplied research confirms code interpreter workflows, but it does not provide a more detailed list of supported programming languages, execution limits, or file-size limits, so those details should not be assumed.
There is an important limitation for developers building agents: function calling is listed as unsupported, and structured outputs are also listed as unsupported. That makes o3-deep-research unsuitable for systems that depend on ordinary function-call schemas or guaranteed JSON-schema-conforming responses. Supported built-in research tools and MCP interfaces are the intended integration path described in the available documentation.
Context and output limits in practice
The 200,000-token context window is useful when a task requires many sources or long documents to be considered together. A legal investigation, technical literature review, or internal-data analysis may need to combine instructions, retrieved passages, prior findings, and supporting evidence. The large window reduces the need to split every investigation into very small independent requests.
The 100,000-token maximum output is similarly suited to detailed reports. However, a maximum is not a recommendation to request the longest possible answer. Very large reports take more time and incur more output-token charges. A better implementation should request the level of detail needed for the decision, organize the response around explicit questions, and preserve source traceability where the application requires it.
Long context also does not eliminate the need for good research design. Relevant sources still need to be selected, contradictory evidence still needs to be examined, and the final response should distinguish retrieved evidence from the model’s interpretation. The model can assist with those steps, but the application remains responsible for review and governance.
Pricing and API access
OpenAI lists o3-deep-research at $10 per 1 million input tokens, $2.50 per 1 million cached input tokens, and $40 per 1 million output tokens. Batch API pricing is also available. These prices are token rates rather than a fixed per-request subscription price, so the total cost depends on the amount of input, cached material, generated output, and any applicable tool usage.
Output is substantially more expensive than uncached input at the listed rates. Requests that produce lengthy research reports can therefore become costly even when the initial prompt is short. Caching may reduce the price of repeated input material when the workflow qualifies for cached-input billing, but the supplied research does not specify the exact eligibility or cache duration rules.
Deep-research requests were released for the Responses API on June 24, 2025. A valid workflow must include at least one supported data source. Web search, remote MCP servers, file search with vector stores, and code interpreter should be treated as components of the request design rather than as general capabilities available in every call.
Strengths and limitations
The strongest case for o3-deep-research is a difficult investigation where the answer depends on collecting and reconciling information from multiple sources. Its large context window supports long source sets, while its research tools can extend the workflow beyond the model’s June 2024 knowledge cutoff. The model can also produce a substantially longer final report than a typical short-answer interaction.
Its principal limitations are equally clear:
- Deprecated status: OpenAI currently lists the model as deprecated, creating lifecycle risk for new deployments.
- Latency: Multi-step research and tool use are inherently less suitable for instant responses than simple generation.
- Cost: Output is listed at $40 per 1 million tokens, and long reports can require substantial output-token usage.
- No ordinary function calling: Agent architectures that expect standard function-call support cannot rely on it.
- No structured outputs: Applications requiring guaranteed JSON-schema output should choose a different approach or model.
- Text-only output: It cannot directly generate images, audio, or video.
- No fine-tuning: The supplied specifications list fine-tuning as unsupported.
These limitations mean that the model’s high research capability should not be confused with universal suitability. A fast, lower-cost model is likely better for routine classification, short answers, or high-volume chat. A model with structured-output and function-calling support is better for tightly controlled business-process automation.
Best use cases
o3-deep-research is a strong fit for work where the quality of the investigation matters more than immediate response speed. Suitable examples include:
- Research reports that compare many web sources.
- Technical literature reviews involving long papers or documentation.
- Legal or policy research that requires collecting and contrasting documents.
- Market research and competitor analysis.
- Scientific or technical investigations supported by calculations.
- Analysis of large internal datasets connected through supported file-search or MCP workflows.
For these tasks, it is useful to define the research question clearly, specify the desired report structure, identify the permitted sources, and decide whether code interpreter is needed. Reviewers should also check important claims against the underlying sources, especially when the output will inform legal, scientific, financial, or operational decisions.
When to choose this model
Choose o3-deep-research when you need a dedicated, tool-assisted investigation with extensive source material, long context, and a detailed textual result. Its combination of web search, file search, MCP connectivity, and code interpreter is more relevant than raw conversational speed when the task involves gathering evidence and explaining how that evidence supports a conclusion.
Choose another type of model when the task is primarily low-latency chat, simple text generation, high-volume processing, image generation, audio or video generation, ordinary function-calling automation, or guaranteed structured JSON. You should also prefer a currently supported successor or alternative when lifecycle stability is more important than reproducing an existing o3-deep-research workflow.
The model’s catalog status makes migration planning especially important. OpenAI explicitly removed the dated o3-deep-research-2025-06-26 snapshot from the API on July 23, 2026, while the available documentation does not provide a separate shutdown date for the canonical alias. Existing users should therefore monitor OpenAI’s deprecation documentation and test the provider’s recommended replacement before a forced migration becomes necessary.
Bottom line
o3-deep-research is a specialized reasoning model for source-heavy investigation, not a general-purpose model for every API task. Its verified advantages are a 200,000-token context window, up to 100,000 output tokens, text and image input, text output, and support for several research-oriented tools through the Responses API. Its cost, likely slower workflow, lack of ordinary function calling and structured outputs, and deprecated status substantially narrow the situations in which it is the right choice.
For existing research systems, the model remains useful to understand and evaluate. For new production work, its deprecation means capability should be weighed against availability and migration risk rather than treated as the only selection criterion.

