Gemini 2.5

Gemini 2.5 Flash

by Google DeepMind · Stable and currently served through the Gemini API with restricted access for users who have actively used Gemini 2.5 models; not deprecated; no shutdown date announced.

Gemini 2.5 Flash is a stable Google DeepMind model for fast, high-volume multimodal processing. It accepts text, images, video, and audio; produces text; and supports configurable thinking, a 1-million-token context window, function calling, grounding, structured outputs, caching, code execution, and batch processing. Its main trade-offs are text-only output, no documented fine-tuning, a January 2025 knowledge cutoff, and restricted current access.

Text Reasoning Coding
Gemini 2.5 Flash is Google DeepMind’s price-performance-focused model for applications that need useful reasoning without the cost or latency of a larger frontier model. It accepts text, images, video, and audio, produces text, and supports a wide range of API features for coding, document processing, agents, data extraction, and large-context analysis.
Outputs

What Gemini 2.5 Flash can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Gemini 2.5
Model type General Purpose
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff January 2025
Release date 2025-06-17
Status Stable and currently served through the Gemini API with restricted access for users who have actively used Gemini 2.5 models; not deprecated; no shutdown date announced.
Knowledge cutoff notes

Google's exact Gemini 2.5 Flash model page lists January 2025 as the knowledge cutoff. Search grounding and other external tools can provide newer information during use but do not change the underlying cutoff.

Model notes

The canonical stable model ID is gemini-2.5-flash. Google lists it as available but restricts access to users who have actively used Gemini 2.5 models, recommending newer models for new projects. The model supports configurable thinking; thinking tokens are included in output billing and count toward the max_output_tokens limit. Thinking can be disabled with a thinking budget of 0 or made dynamic with -1, with a documented maximum thinking budget of 24,576 tokens. The model accepts text, images, video, and audio but natively outputs text only. Google documents structured outputs, but this record leaves legacy JSON mode unknown because structured-output support is not treated as automatic proof of a separate JSON-mode capability. The dated preview model gemini-2.5-flash-preview-09-2025 is a separate retired preview identifier and should not be treated as the stable model.

Cost

Model pricing

Input $0.30 per 1M text/image/video tokens; $1.00 per 1M audio tokens; batch input $0.15 per 1M text/image/video tokens.
Output $2.50 per 1M tokens, including thinking tokens.
Model guide

Gemini 2.5 Flash: Google’s Fast, Cost-Efficient Reasoning Model

Gemini 2.5 Flash is a stable Google DeepMind reasoning model designed for low-latency, high-volume workloads. It combines multimodal input, configurable thinking, a 1-million-token context window, tool use, structured outputs, grounding, caching, and batch processing at lower standard token prices than many more capability-focused models.

What is Gemini 2.5 Flash?

Gemini 2.5 Flash is a stable multimodal reasoning model from Google DeepMind. It is designed for production workloads where response speed, operating cost, and scalable processing matter alongside reasoning ability. Typical uses include high-volume classification, document and media analysis, coding assistance, tool-using agents, and applications that need to process very large inputs.

The model’s canonical Gemini API identifier is gemini-2.5-flash. It was released on June 17, 2025, and remains listed as available rather than deprecated. However, Google currently restricts access to users who have actively used Gemini 2.5 models and recommends newer model lines for new projects. Google has not announced a shutdown date for the stable model.

Gemini 2.5 Flash belongs to Google’s Gemini model family, but this page focuses specifically on the stable 2.5 Flash model rather than the broader Gemini catalog. Its central positioning is straightforward: deliver reasoning and multimodal capabilities at a speed and price suitable for large-scale use.

Main capabilities and supported modalities

Gemini 2.5 Flash accepts four input types: text, images, video, and audio. Its native output is text only. In practical terms, it can inspect an image, summarize a video, transcribe or analyze audio, or combine those inputs with written instructions, but it does not natively generate images, videos, or audio.

  • Text, image, video, and audio input
  • Text output
  • Configurable reasoning, called thinking in the Gemini API
  • Function calling and external tool use
  • Google Search and Google Maps grounding
  • Code execution, file search, and URL context
  • Structured machine-readable output
  • Streaming responses
  • Context caching
  • Batch, flex, and priority inference options

Structured output lets an application request responses that follow a defined schema, which is useful for tasks such as extracting fields from invoices, classifying support messages, or converting unstructured documents into records. This documented structured-output capability should not automatically be treated as proof of a separate legacy JSON-mode feature; the supplied model record leaves that distinction unknown.

Reasoning, context, and output limits

Gemini 2.5 Flash supports configurable thinking. Applications can disable thinking with a thinking budget of zero, use dynamic thinking, or set a maximum thinking budget. The documented maximum thinking budget is 24,576 tokens. Thinking tokens are part of the model’s output accounting and use capacity that would otherwise be available for visible answer text.

The model has an input context limit of 1,048,576 tokens, or just over one million tokens. A context window is the amount of information the model can consider in one request, so this limit is particularly relevant to long documents, large code repositories, collections of files, and extended multimodal inputs.

The maximum output limit is 65,536 tokens. For requests that use thinking, the configured maximum output limit includes both internal thinking tokens and visible response tokens. Setting the limit too low can therefore reduce the space available for the answer or cause a response to be truncated.

SpecificationVerified detail
Model IDgemini-2.5-flash
Release dateJune 17, 2025
Context length1,048,576 tokens
Maximum output65,536 tokens
Knowledge cutoffJanuary 2025
Native outputText
Fine-tuningNot supported

Pricing and cost trade-offs

At the standard paid tier, Gemini 2.5 Flash costs $0.30 per 1 million input tokens for text, image, and video content. Audio input is priced higher at $1.00 per 1 million tokens. Output costs $2.50 per 1 million tokens, including thinking tokens.

Batch processing reduces the input price for text, image, and video to $0.15 per 1 million tokens. Batch processing is intended for workloads that do not require an immediate response, such as overnight document classification or large-scale data transformation. Google also provides flex and priority inference options, although the applicable service terms and prices can vary.

Context caching costs $0.03 per 1 million cached text, image, or video tokens and $0.10 per 1 million cached audio tokens. Cached content also has a storage charge of $1.00 per 1 million tokens per hour. Caching can be useful when an application repeatedly sends the same long instructions, reference material, or documents.

Google Search grounding and Google Maps grounding have separate request-based charges after applicable free allowances. Free-tier access, quotas, and rate limits may differ from paid usage. The prices above describe the standard paid token rates supplied for the model, not a guarantee of total application cost.

Tools and API support

Gemini 2.5 Flash supports function calling, which allows the model to request an operation from an application rather than attempting to perform that operation itself. For example, an application could expose a database lookup, inventory check, or calendar function and then decide whether to execute the requested call.

The model also supports Google Search grounding and Google Maps grounding. Grounding connects a response to external information or location data, which can be more useful for current or geographically specific questions than relying only on the model’s January 2025 knowledge cutoff. Grounding does not change the underlying cutoff; it supplies additional information during a request.

Other documented tools and features include code execution, file search, URL context, streaming, structured outputs, context caching, and batch processing. These features make the model suitable for agentic applications, where the model interprets a task, uses tools, and returns a structured result.

Coding and data-processing suitability

Gemini 2.5 Flash is suited to coding assistance, code analysis, data extraction, and transformation tasks. Its large context window can help when an application needs to provide extensive source code, technical documentation, or multiple related files in one request. Code execution and structured outputs can further support workflows that inspect data and return consistent records.

Its value for coding is primarily operational rather than a claim of maximum coding performance. A fast, lower-cost model can be preferable for generating routine code snippets, reviewing many files, classifying errors, or handling repeated development tasks. More demanding projects may benefit from a newer or more capability-focused model, particularly if maximum reasoning quality matters more than throughput and price.

Strengths and limitations

Where Gemini 2.5 Flash is strong

  • Large context: The 1,048,576-token input limit supports very long documents, codebases, and multimodal workloads.
  • Balanced cost and speed: Its standard input price is low enough for high-volume processing, while the model is positioned for low-latency use.
  • Multimodal understanding: It can process text, images, video, and audio in addition to ordinary text prompts.
  • Configurable reasoning: Developers can trade reasoning depth against latency and token consumption by adjusting thinking.
  • Production features: Function calling, grounding, structured outputs, caching, streaming, code execution, and batch processing support application development beyond simple chat.

Important limitations

  • Text-only output: The model does not natively generate images, video, or audio.
  • Knowledge cutoff: Its underlying knowledge cutoff is January 2025, so current information may require search or another external tool.
  • No fine-tuning: The supplied documentation does not list fine-tuning support for this model.
  • Access restrictions: Current Gemini 2.5 access is limited to users who have actively used Gemini 2.5 models, and Google directs new projects toward newer model lines.
  • Thinking affects capacity and billing: Thinking tokens count toward output usage and the maximum output limit.
  • Preview confusion: Preview identifiers such as gemini-2.5-flash-preview-09-2025 are separate models and should not be confused with the stable gemini-2.5-flash identifier.

When to choose Gemini 2.5 Flash

Choose Gemini 2.5 Flash when the application needs a combination of multimodal input, useful reasoning, fast responses, and predictable high-volume economics. It is a practical fit for document extraction, media question answering, coding support, classification, large-context summarization, and agents that call tools or use external information.

It is especially attractive when the same long context is reused repeatedly, because context caching may reduce the cost of resending that material. Batch processing is a good fit for large queues where immediate responses are unnecessary. Configurable thinking also lets a developer use less reasoning for routine requests and more for difficult ones.

Another model type may be more appropriate when the application needs native image, video, or audio generation, requires fine-tuning, or prioritizes the newest available capabilities over compatibility with Gemini 2.5 Flash. A more capability-focused reasoning model may also be preferable for the hardest tasks if higher latency and cost are acceptable. Conversely, a simpler non-reasoning model may be a better choice for extremely basic, latency-sensitive operations where multimodal understanding and deeper reasoning are unnecessary.

Bottom line

Gemini 2.5 Flash is a stable, text-output reasoning model built for scale. Its defining practical combination is a million-token context window, multimodal input, adjustable thinking, broad tool support, and relatively low standard token pricing. Those characteristics make it useful for production pipelines and agents that need to process substantial information quickly.

Its limitations are equally important: it does not generate media, it cannot be fine-tuned according to the supplied documentation, its underlying knowledge ends in January 2025, and current access is restricted while Google recommends newer models for new projects. For existing Gemini 2.5 users and workloads that value speed, cost control, and large-context multimodal processing, it remains a relevant option.


Answers to Frequently Asked Questions

What are the main limitations of Gemini 2.5 Flash?
Gemini 2.5 Flash produces text only, has a January 2025 knowledge cutoff, does not support fine-tuning according to the supplied documentation, and has current access restrictions. Google recommends newer model lines for new projects.
How much does Gemini 2.5 Flash cost?
At the standard paid tier, text, image, and video input costs $0.30 per 1 million tokens, audio input costs $1.00 per 1 million tokens, and output costs $2.50 per 1 million tokens, including thinking tokens. Batch processing and context caching offer lower input rates for eligible workloads.
What are the context and output limits of Gemini 2.5 Flash?
The model supports an input context of up to 1,048,576 tokens and a maximum output of 65,536 tokens. When thinking is enabled, thinking tokens count toward the output limit and usage.
What is Gemini 2.5 Flash?
Gemini 2.5 Flash is a stable multimodal reasoning model from Google DeepMind designed for fast, cost-efficient production workloads. Its canonical Gemini API identifier is `gemini-2.5-flash`.
What types of input and output does Gemini 2.5 Flash support?
Gemini 2.5 Flash accepts text, images, video, and audio as input. Its native output is text only, so it does not directly generate images, video, or audio.


Sources 7
Provider

About Google DeepMind