What is Gemini Deep Research?
Gemini Deep Research is a preview agent from Google DeepMind for carrying out multi-step investigations. Its model identifier is deep-research-preview-04-2026. Instead of simply generating an answer from one prompt, it manages a research process: it can plan the work, search for information, retrieve and read sources, analyze material, and then write a report with citations.
This makes it closer to a managed research workflow than to a conventional chat model. A request may involve multiple searches and tool calls before a final answer is produced. The result is intended for questions where breadth, source comparison, and synthesis matter more than an immediate response.
Google lists the model as a preview and provides access through the Gemini Interactions API in Google AI Studio and the Gemini API. It is not available through the standard generate_content endpoint.
How the research agent works
Deep Research runs as a background interaction because a substantial investigation can take several minutes. Background execution allows the agent to continue working while it searches, reads and reasons, rather than requiring the client application to hold a normal synchronous request open.
The workflow can include collaborative planning, in which the research plan is prepared before execution. During the investigation, the agent can use Google Search and URL Context to find and retrieve online material. It can also use Code Execution for calculations and data analysis. File Search and remote MCP servers are available for supported integrations, giving applications additional ways to supply documents or connect external tools.
Streaming progress and output are supported, which is useful for interfaces that need to show that a long-running task is still active. However, streaming does not make the model a low-latency chat endpoint: the complete research task can still take considerably longer than a normal model response.
Inputs, outputs and technical limits
The model accepts text, images, PDFs, audio and video as inputs. Its main output is text: a research report that can include citations and synthesized findings. When visualization support is enabled, it can also generate images or visualizations such as charts. These visual outputs are an additional research capability, not a replacement for a general-purpose image-generation product.
| Specification | Reported detail |
|---|---|
| Model ID | deep-research-preview-04-2026 |
| Status | Preview |
| Context window | 1,048,576 tokens |
| Maximum output | 65,536 tokens |
| Primary interface | Interactions API |
| Execution style | Background, asynchronous interaction |
| Structured outputs | Not currently supported |
The one-million-token context window provides substantial room for source material and intermediate information, while the 65,536-token output limit allows for long reports. These limits do not guarantee that every task will use the full capacity: actual work depends on the question, the sources found, tool activity and the agent's research decisions.
Tools, reasoning and coding
Deep Research's main distinction is its ability to combine reasoning with a sequence of research actions. It can break a broad question into smaller steps, decide what information to seek, compare sources, and combine findings into a coherent answer. This is useful when the research path cannot be known in advance.
Its documented tools include Google Search, URL Context and Code Execution. Search helps locate current web information, while URL Context allows the agent to work directly with retrieved pages. Code Execution supports calculations and data analysis, such as processing numerical information or producing a chart from research findings. File Search and remote MCP server support can extend the workflow to supplied files and external services.
These tools improve the agent's ability to investigate current or source-heavy questions, but they do not remove the need for review. Search results can be incomplete, sources can disagree, and a citation can be technically present without making the underlying claim reliable. Important reports should therefore be checked against the linked sources, especially when decisions depend on accuracy.
The supplied evaluation rates its reasoning at 8 out of 10 and coding at 7 out of 10. Those are editorial scores for comparison, not scores published by Google and not benchmark results. The provider's documented capability is that the agent can perform multi-step reasoning and use code execution during research.
Pricing and cost trade-offs
Deep Research uses variable pay-as-you-go pricing rather than a fixed subscription price or a single universal per-request fee. Google states that cost depends on factors including searching, source reading, reasoning, tool use and token processing. The supplied documentation estimates approximately $1 to $3 for a typical standard Deep Research task, but the actual amount can vary with research complexity and preview pricing.
This pricing model changes how the agent should be used. A simple factual question may not justify launching a multi-step investigation. A market review, due-diligence assignment or literature survey may justify the cost because the value comes from collecting and synthesizing many sources rather than producing a quick paragraph.
The supplied evaluation gives the model a cost score of 6 out of 10 and a speed score of 8 out of 10. These are editorial assessments, not provider-published ratings. In practical terms, Deep Research is faster and more efficient than a hypothetical very long-running research agent, but it remains slower and more expensive than ordinary synchronous chat for small tasks.
Main strengths and limitations
Where it is strongest
- Multi-step investigation: it can plan and execute a research process instead of relying on one model completion.
- Source-oriented answers: it is designed to synthesize information into reports with citations.
- Current information: Google Search and URL Context can bring web sources into an individual task.
- Analysis of varied material: text, images, PDFs, audio and video can be supplied as inputs.
- Data and visualization support: Code Execution can assist with calculations and analysis, while visualization support can produce charts or other visual outputs.
- Large working context: the documented 1,048,576-token context window can accommodate substantial source material.
What it does not solve
- It is not synchronous: background execution and potentially long runtimes make it unsuitable for interactive, immediate-response applications.
- It is still a preview: behavior, pricing, limits and availability may change.
- Costs are variable: a research task's final cost depends on how much searching, reading, reasoning and tool use it requires.
- Structured outputs are unavailable: Google documents that structured outputs are not currently supported, making the model a poor fit for pipelines that require strict schema-conforming JSON.
- Citations need checking: source coverage and citations do not guarantee that every conclusion is correct or that every source is authoritative.
- It is not ideal for simple prompts: using a research agent for a short answer can add unnecessary delay and expense.
When to choose Gemini Deep Research
Choose Gemini Deep Research when the task requires several stages of discovery and synthesis. Suitable examples include market analysis, competitive research, due diligence, literature reviews, policy research, technology assessments and reports that must show where claims came from. It is especially useful when the researcher wants the system to discover relevant sources rather than supplying a complete, known set of documents in advance.
For example, a request to compare competitors across pricing, product capabilities, recent announcements and customer segments may require many searches and source checks. Deep Research can organize that investigation and produce a cited synthesis. A literature review can similarly benefit from collecting information across multiple documents and identifying points of agreement or disagreement.
Another option may be more appropriate when the task needs a fast answer, a predictable schema, a low fixed cost or a synchronous API response. A conventional chat model is generally better for rewriting, summarization of a known document, short explanations and routine coding assistance. A standard model with structured-output support is preferable when downstream software must parse the response reliably. Deep Research Max is a separate preview model positioned for more comprehensive and longer-running investigations; it should not be treated as the same model or as part of this model's specifications.
How to evaluate a Deep Research result
Because the agent performs autonomous research, evaluation should cover both the final writing and the research process. Check whether the cited sources actually support the claims, whether primary sources were used where available, and whether important opposing evidence was considered. Review dates carefully when the subject changes quickly, and distinguish facts taken directly from sources from the agent's interpretation.
For production use, applications should also track task duration, tool activity and cost. The lack of structured outputs means that report extraction may require additional processing, and the asynchronous interaction model requires application logic for job status, progress and completion. These operational requirements are central to deciding whether the model fits a workflow.
Bottom line
Gemini Deep Research is a preview agent for source-heavy investigations rather than a general-purpose chat endpoint. Its value lies in combining planning, web research, document and data analysis, tool use and cited report generation in one managed workflow. The trade-off is meaningful: tasks can take minutes, costs vary with research depth, and the model does not currently provide structured outputs. For complex investigations where source breadth and synthesis matter more than immediate responses, it is a compelling specialized option; for short, deterministic or schema-driven tasks, a conventional model is usually the better choice.

