What is Gemini 3.1 Pro?
Gemini 3.1 Pro is Google DeepMind’s third-generation Pro model and is currently offered as a preview model. It is designed for demanding tasks where a short, simple response is not enough: complex analysis, advanced coding, long-document review, multimodal research, and applications that need several steps of planning and tool use.
The canonical API identifier is gemini-3.1-pro-preview. Developers can access it through Google AI Studio and the Gemini API, with additional availability through Google Cloud Vertex AI and selected Google developer and enterprise products. A related gemini-3.1-pro-preview-customtools endpoint is optimized for workflows using custom tools and bash commands, but it should be understood as a related endpoint rather than a separate base model.
Within Google’s current model lineup, Gemini 3.1 Pro occupies a high-capability Pro position rather than a low-cost, high-throughput role. Its purpose is to handle difficult inputs and multi-step work, even when that requires more computation, more latency, and a higher per-token cost.
Supported inputs and outputs
Gemini 3.1 Pro accepts five main input types: text, images, video, audio, and PDF documents. This allows a single request to combine, for example, a written research question with pages from a PDF, screenshots, a recorded meeting, or a video clip.
The model’s output is text. It can explain findings, produce code, return structured responses, and describe decisions made during a tool-assisted workflow, but it does not natively generate images, video, or audio. This distinction matters when selecting it for a multimodal application: Gemini 3.1 Pro is multimodal in what it can understand, not in the kinds of media it directly creates.
| Specification | Verified value |
|---|---|
| Provider | Google DeepMind |
| Model identifier | gemini-3.1-pro-preview |
| Status | Preview |
| Input modalities | Text, image, video, audio, and PDF |
| Maximum input context | 1,048,576 tokens |
| Maximum output | 65,536 tokens |
| Output modality | Text |
A token is a small unit of text or other encoded input. The exact number of tokens a document uses depends on its content and format, but the one-million-token context limit is intended for unusually large inputs such as extensive repositories, long reports, or collections of multimodal materials.
Reasoning and coding capabilities
Google positions Gemini 3.1 Pro for advanced reasoning rather than only conversational question answering. It is intended to work through complicated problems, compare evidence across a large context, plan multiple steps, and use tools when an answer depends on external information or computation. Thinking tokens are included in the billed output-token pricing, so deeper reasoning can affect cost even when the visible answer is relatively short.
For software work, the model is aimed at advanced engineering tasks such as analyzing unfamiliar codebases, tracing behavior across many files, proposing implementation changes, reviewing technical designs, and helping coordinate multi-step development workflows. Its large context can be useful when an answer depends on relationships spread across a repository rather than on one isolated code snippet.
These capabilities should not be confused with a guarantee of correctness. The supplied research gives editorial scores of 9 out of 10 for reasoning and coding, but those are comparative editorial assessments, not Google-published benchmark scores. The model can still produce incorrect explanations, flawed code, or an unsuitable plan, especially when the input is ambiguous or a tool returns incomplete information.
Tools, grounding, and structured responses
Gemini 3.1 Pro supports function calling, which lets an application expose defined operations that the model can request. For example, an application might allow the model to query a database, call an internal service, or start a controlled software workflow. The application remains responsible for executing and authorizing those functions.
It also supports code execution, URL context, Google Search grounding, and Google Maps grounding. Grounding connects an answer to retrieved information or a specified source instead of relying solely on the model’s internal generation. Code execution can help with calculations or data-processing steps, while URL context can make a supplied web resource part of the model’s working information.
Structured outputs are supported, allowing developers to request responses that follow a defined schema. This can make it easier to pass model results into software. However, the available documentation does not independently establish a separate legacy “JSON mode” capability for this model. Structured outputs should therefore not automatically be described as a distinct JSON-mode feature.
Caching is available for repeated large prompts, and batch processing is available for workloads that can be handled asynchronously. Google also documents flex and priority consumption options where applicable. File Search is supported in Google AI Studio only according to the supplied model notes.
Context window and output limits
The maximum input context is 1,048,576 tokens, while the maximum output is 65,536 tokens. The large input limit is particularly relevant when the model must compare many sources, inspect a substantial codebase, or maintain context across a long agentic task.
A large context window does not mean every long prompt will receive equally careful treatment. Very large inputs can still increase processing time and cost, and the usefulness of the result depends on how clearly the task is framed. For practical applications, it is still sensible to supply relevant material, identify the desired outcome, and ask the model to distinguish evidence from assumptions.
Gemini 3.1 Pro pricing
Gemini 3.1 Pro Preview uses token-based pricing. For prompts of 200,000 tokens or fewer, standard paid pricing is $2.00 per million input tokens and $12.00 per million output tokens, including thinking tokens. For prompts above 200,000 tokens, the price rises to $4.00 per million input tokens and $18.00 per million output tokens.
| Prompt size | Input price | Output price |
|---|---|---|
| Up to 200,000 tokens | $2.00 per million tokens | $12.00 per million tokens |
| More than 200,000 tokens | $4.00 per million tokens | $18.00 per million tokens |
Context caching is priced at $0.20 per million tokens for prompts up to 200,000 tokens and $0.40 per million tokens for larger prompts, with separate storage charges. Batch, flex, and priority options may have their own pricing or availability conditions. The token prices above are usage prices, not a recurring consumer subscription fee.
The practical trade-off is clear: Gemini 3.1 Pro can be economical when its reasoning and context capacity prevent costly application-side processing, but it is not positioned as the cheapest choice for simple prompts or very high-volume generation. Caching can help when the same large context is reused, while smaller or faster model options may be more appropriate for routine classification, short answers, or latency-sensitive interactions.
Main strengths and limitations
Strengths
- Very large input context for long documents, repositories, and multimodal collections.
- Broad multimodal understanding across text, images, video, audio, and PDFs.
- Support for complex reasoning, advanced coding, and multi-step agentic workflows.
- Function calling, code execution, search and Maps grounding, URL context, and structured outputs.
- Support for caching and batch processing in suitable developer workflows.
Limitations
- The model is still in preview, so behavior, availability, pricing, and endpoints may change.
- It produces text only and does not natively generate images, video, or audio.
- Large prompts and long reasoning can make usage more expensive and potentially slower than simpler alternatives.
- Tool support does not remove the need for application-side permissions, validation, monitoring, and error handling.
- There is no authoritative knowledge-cutoff date identified in the supplied documentation.
- Model outputs can be inaccurate or inconsistent, including in code and grounded workflows.
Best use cases
Gemini 3.1 Pro is a strong fit when the task combines difficult reasoning with substantial context or external tools. Suitable examples include:
- Reviewing large technical specifications, legal-style documents, research collections, or PDF archives.
- Analyzing a large software repository and planning changes across multiple files.
- Building coding agents that need function calling, code execution, or custom tools.
- Combining written questions with images, recordings, video, or documents during research.
- Creating grounded applications that use Google Search, Google Maps, URLs, or application-defined functions.
- Generating structured text results for downstream software systems.
When to choose Gemini 3.1 Pro
Choose Gemini 3.1 Pro when the main challenge is complexity: the model must reason through a difficult problem, inspect a large context, understand several media types, or coordinate tools. Its value is greatest when those capabilities replace manual investigation or multiple separate processing steps.
Consider another option when the task is simple, cost-sensitive, or strongly latency-sensitive. A smaller or faster model may be preferable for short classification tasks, repetitive extraction, basic chat, and high-volume generation. A dedicated image, video, or audio generation model is more appropriate when the required output is media rather than text. A stable generally available model may also be a better operational choice when preview-status changes would create unacceptable risk.
Overall, Gemini 3.1 Pro is best understood as a high-capability preview model for difficult, context-heavy, tool-assisted work. Its one-million-token input capacity and broad tool support are its clearest practical differentiators, while preview status, text-only output, and token costs are the main reasons not to use it for every workload.

