Gemini 2.5

Gemini 2.5 Pro

by Google DeepMind · Stable and generally available through the Gemini API; access is currently limited to users who have actively used Gemini 2.5 models. Google states that the model is not deprecated and will continue to be served until further notice.

Gemini 2.5 Pro is a stable Google DeepMind reasoning model for advanced coding, mathematics, STEM analysis, multimodal understanding, and long-context workloads. It accepts text, images, audio, video, and PDFs, supports a 1,048,576-token context window, and offers tool use, structured outputs, code execution, file search, and grounding features. Its main trade-offs are text-only output, higher cost and latency than speed-focused models, a January 2025 knowledge cutoff, and restricted access for new users.

Text Reasoning Coding
Gemini 2.5 Pro is Google DeepMind’s high-end stable model for demanding reasoning and coding tasks. It combines adaptive thinking, native multimodal input, long-context processing, tool use, structured outputs, code execution, file search, URL context, and Google Search and Maps grounding. Its main advantage is depth and context capacity rather than minimum latency or lowest cost.
Outputs

What Gemini 2.5 Pro can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
6/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Gemini 2.5
Model type Reasoning
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff January 2025
Release date 2025-06-17
Status Stable and generally available through the Gemini API; access is currently limited to users who have actively used Gemini 2.5 models. Google states that the model is not deprecated and will continue to be served until further notice.
Knowledge cutoff notes

Google’s current model documentation and model card identify January 2025 as the knowledge cutoff. Search grounding and other external-context features can provide newer information during use but do not change the underlying cutoff.

Model notes

Gemini 2.5 Pro is a sparse mixture-of-experts thinking model with native multimodal input and text-only output. The exact stable model ID is gemini-2.5-pro. It supports code execution, file search, function calling, Google Search grounding, Google Maps grounding, URL context, structured outputs, caching, Batch API, Flex inference, and Priority inference. Standard context caching is $0.125 per 1M tokens for prompts up to 200K and $0.25 per 1M tokens above 200K, with storage at $4.50 per 1M tokens per hour. Google Search grounding and Google Maps grounding are separately priced tools. Fine-tuning is not supported for current Gemini models. The model’s January 2025 knowledge cutoff is distinct from any information supplied through search, retrieval, uploaded files, or other external context. Scores are editorial comparative estimates, not vendor ratings.

Cost

Model pricing

Input $1.25 per 1M tokens for prompts up to 200K tokens; $2.50 per 1M tokens for prompts over 200K tokens. Batch/Flex rates are $0.625/$1.25 respectively.
Output $10.00 per 1M tokens for prompts up to 200K tokens; $15.00 per 1M tokens for prompts over 200K tokens. Batch/Flex rates are $5.00/$7.50 respectively. Prices include thinking tokens.
Model guide

Gemini 2.5 Pro: Google’s Long-Context Reasoning Model for Complex Coding

Gemini 2.5 Pro is Google DeepMind’s stable reasoning model for advanced coding, mathematics, STEM analysis, multimodal understanding, and long-context workloads. It accepts text, images, audio, video, and PDF inputs, supports a 1,048,576-token context window, and produces text outputs of up to 65,536 tokens.

What is Gemini 2.5 Pro?

Gemini 2.5 Pro is a stable reasoning model from Google DeepMind’s Gemini family. It is designed for tasks where the system must analyze substantial information, follow multiple steps, or produce technically detailed results. Typical examples include advanced software development, repository analysis, difficult mathematics, scientific and technical research, long-document question answering, and multimodal analysis.

The stable model identifier is gemini-2.5-pro. It became generally available on June 17, 2025, through the Gemini API. As of September 2026, Google continues to serve it, but access is limited to users who have actively used Gemini 2.5 models. Google recommends newer models for new projects while stating that Gemini 2.5 Pro is not deprecated and will continue to be served until further notice.

This makes Gemini 2.5 Pro particularly relevant to existing applications that depend on its behavior or context capacity. New projects should evaluate the current Google model catalog before committing to it, especially if unrestricted access to the newest models is important.

Where it fits in Google’s model lineup

Gemini 2.5 Pro occupies the advanced, reasoning-oriented position within the Gemini 2.5 family. It is not primarily optimized for the fastest possible responses or the lowest token cost. Instead, it is intended for complex work where answer depth, coding ability, long-context processing, and multimodal understanding are more important than minimal latency.

In practical terms, it sits closer to a general-purpose expert model than to a lightweight model used for high-volume, simple requests. The model can handle ordinary text prompts, but its value is clearest when a task involves large inputs, difficult reasoning, multiple information sources, or tool-assisted workflows.

Inputs, outputs, and context limits

Gemini 2.5 Pro accepts several input types in the same model interface:

  • Text
  • Images
  • Audio
  • Video
  • PDF documents

Its maximum context window is 1,048,576 tokens. A context window is the amount of information the model can consider in one request and its associated conversation or task state. A one-million-token limit is useful for large codebases, extensive documentation, long research material, and collections of files that would need to be split across requests with a smaller model.

The maximum output is 65,536 tokens. The model’s output is text only: it does not directly generate images, audio, or video. Its ability to analyze images, audio, video, and PDFs should therefore not be confused with native media generation.

SpecificationGemini 2.5 Pro
Stable model IDgemini-2.5-pro
Input modalitiesText, images, audio, video, and PDF
Output modalityText
Context window1,048,576 tokens
Maximum output65,536 tokens
Knowledge cutoffJanuary 2025
Release dateJune 17, 2025

Reasoning and coding capabilities

Gemini 2.5 Pro is a thinking model: it can perform internal reasoning before presenting its answer. This is useful for multi-step tasks in which a quick pattern match is less reliable than a structured analysis. Examples include tracing a software bug through several files, comparing alternative technical designs, working through advanced mathematics, or extracting conclusions from a long technical document.

Its coding uses include code generation, code review, repository-scale understanding, refactoring guidance, debugging, and software-development agents that call tools. The large context window allows developers to provide more of a project’s source code, documentation, configuration, and test output in one task rather than summarizing everything manually.

Google’s model card reports strong results across reasoning, coding, visual reasoning, video understanding, and long-context evaluations. These are provider-reported evaluation results rather than a guarantee of performance on every application. Real-world accuracy can still vary with prompt quality, input complexity, tool configuration, and the need for current information.

Tools, grounding, and structured output

Gemini 2.5 Pro supports function calling, which allows an application to give the model access to defined external operations. For example, a program can let the model request a database lookup, invoke a business function, or start a controlled software action. The model does not automatically gain unrestricted access to those systems; the application decides which functions exist and whether to execute each requested call.

Supported tools and related features include code execution, file search, URL context, Google Search grounding, Google Maps grounding, caching, Batch inference, Flex inference, and Priority inference. Code execution can help with tasks that benefit from running calculations or analyzing data. File search and URL context can provide material outside the model’s built-in knowledge. Search and Maps grounding are separate services with their own usage charges.

The model also supports structured outputs using supported JSON Schema features. Structured output is useful when an application needs predictable fields, such as extracting names, dates, classifications, or code-review findings. It should not be treated as a guarantee that every response is semantically correct: the response can follow the requested structure while still containing an inaccurate conclusion.

Gemini 2.5 Pro API pricing

Standard Gemini API pricing depends on prompt length. For prompts of 200,000 tokens or fewer, input costs $1.25 per 1 million tokens. Output costs $10 per 1 million tokens, including thinking tokens. For prompts above 200,000 tokens, input costs $2.50 per 1 million tokens and output costs $15 per 1 million tokens.

Standard API usageUp to 200,000 input tokensOver 200,000 input tokens
Input$1.25 per 1M tokens$2.50 per 1M tokens
Output, including thinking tokens$10 per 1M tokens$15 per 1M tokens

Batch and Flex inference use lower rates: $0.625 per 1 million input tokens and $5 per 1 million output tokens for prompts up to 200,000 tokens, rising to $1.25 input and $7.50 output for larger prompts. Priority inference costs more than the standard option.

Context caching is priced at $0.125 per 1 million tokens for prompts up to 200,000 tokens and $0.25 per 1 million tokens for larger prompts. Cached-content storage costs $4.50 per 1 million tokens per hour. Google Search grounding includes 1,500 requests per day at no additional charge on the paid tier, followed by $35 per 1,000 grounded prompts. Google Maps grounding includes 10,000 requests per day at no additional charge on the paid tier, followed by $25 per 1,000 grounded prompts.

Actual spending depends on both input and output volume. Long prompts can also move a request into the higher pricing tier, so the model’s large context window should be used because the task benefits from it, not simply because the capacity is available.

Main strengths and limitations

The most important strength of Gemini 2.5 Pro is the combination of advanced reasoning, coding ability, multimodal input, and a one-million-token context window. It can analyze large heterogeneous inputs and connect information across text, images, audio, video, and PDF documents. Tool support extends it beyond a standalone question-and-answer system and makes it suitable for controlled agentic workflows.

Its main limitations are equally important:

  • It produces text rather than native image, audio, or video output.
  • Its January 2025 knowledge cutoff means built-in knowledge does not automatically include later events.
  • Search grounding, URL context, uploaded files, and other external sources can provide newer information, but they do not change the underlying cutoff.
  • Thinking and long responses can make it slower or more expensive than a smaller, speed-focused model.
  • Access is currently limited to users who have actively used Gemini 2.5 models.
  • Google recommends newer models for new projects, so long-term availability and migration planning should be considered.
  • Responses can still be inaccurate, even when the model presents a detailed chain of reasoning or uses external tools.

Speed and cost trade-offs

Gemini 2.5 Pro is best viewed as a quality-and-capability choice rather than a low-latency choice. Its reasoning process, large context support, and high output ceiling are valuable for difficult tasks, but they can increase response time and token consumption. Output pricing includes thinking tokens, so a request that requires substantial internal reasoning may cost more than a short direct answer suggests.

For simple classification, short summaries, routine extraction, or very high-volume requests, a smaller or faster model may be more economical. For a complex code review or long-document investigation, paying more for Gemini 2.5 Pro can be justified if reducing manual preparation or improving answer depth matters more than response speed.

When to choose Gemini 2.5 Pro

Choose Gemini 2.5 Pro when the task benefits from several of the following characteristics:

  • Advanced software development, debugging, or code review
  • Analysis of large repositories or long technical documents
  • Complex mathematics, science, or engineering questions
  • Multimodal analysis involving documents, images, audio, or video
  • Long-context question answering across extensive source material
  • Tool-using agents that need function calling or code execution
  • Structured extraction from large or mixed-format inputs

It is less appropriate when the primary requirement is native media generation, minimum latency, the lowest possible cost, or unrestricted access to Google’s newest model line. It is also a poor fit for applications that need current facts without configuring search, retrieval, URL context, or another external information source.

Availability and practical verdict

Gemini 2.5 Pro remains a stable Google DeepMind model available through the Gemini API, with no announced deprecation or shutdown date in the supplied documentation. However, the restriction to users who have actively used Gemini 2.5 models makes availability a practical consideration, particularly for new deployments.

For existing systems, it remains a capable choice when its behavior, multimodal input, tool support, or million-token context window solves a specific problem. For new systems, compare it with newer Google models and test representative workloads before adoption. The clearest reason to choose Gemini 2.5 Pro is the combination of deep reasoning, advanced coding, broad input support, and very large context—not image or video generation, ultra-low latency, or the lowest per-request cost.


Answers to Frequently Asked Questions

What are the main limitations of Gemini 2.5 Pro?
Gemini 2.5 Pro has a January 2025 knowledge cutoff, produces text rather than native media, can be slower and more expensive than smaller models, and may still generate inaccurate responses. Access is also limited to users who have actively used Gemini 2.5 models, and Google recommends newer models for new projects.
How much does Gemini 2.5 Pro cost through the Gemini API?
For prompts up to 200,000 input tokens, standard pricing is $1.25 per 1 million input tokens and $10 per 1 million output tokens, including thinking tokens. For prompts above 200,000 tokens, pricing rises to $2.50 per 1 million input tokens and $15 per 1 million output tokens.
What is Gemini 2.5 Pro used for?
Gemini 2.5 Pro is designed for complex tasks such as advanced software development, repository analysis, debugging, difficult mathematics, scientific research, long-document question answering, and multimodal analysis.
What is the context window and maximum output of Gemini 2.5 Pro?
Gemini 2.5 Pro has a maximum context window of 1,048,576 tokens and supports outputs of up to 65,536 tokens. It accepts text, images, audio, video, and PDF inputs, but its output is text only.


Sources 7
Provider

About Google DeepMind