Gemini 3

Gemini 3.6 Flash

by Google DeepMind · Stable; previous-generation Flash model; currently accessible; no shutdown date announced

Gemini 3.6 Flash is a stable, previous-generation Google model optimized for fast multimodal understanding, coding, long-context analysis, and tool-using agents. It accepts text, image, audio, video, and PDF inputs, supports up to 1,048,576 input tokens and 65,536 output tokens, and generates text rather than images, video, or audio.

Text Reasoning Coding
Gemini 3.6 Flash is a stable Google Gemini 3 model designed to balance speed, multimodal understanding, coding ability, and cost. It is no longer the newest Flash model because Gemini 3.7 Flash and Gemini 3.8 Flash are available, but it remains accessible through the Gemini API and related Google platforms. Its combination of a one-million-token context window, broad input support, and fast tool-oriented execution makes it suitable for practical production workloads.
Outputs

What Gemini 3.6 Flash can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3
Model type General Purpose
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff March 2026
Release date 2026-07-21
Status Stable; previous-generation Flash model; currently accessible; no shutdown date announced
Knowledge cutoff notes

Google's model card identifies March 2026 as the knowledge cutoff, while noting that information coverage varies by domain and may be limited to January 2025 for some areas. Search grounding and external tools can provide newer information during use but do not change the underlying cutoff.

Model notes

Canonical stable model identifier is gemini-3.6-flash. Google describes the model as a workhorse Gemini 3 model with improved coding, knowledge work, multimodal performance, and token efficiency compared with Gemini 3.5 Flash. Supported inputs include text, image, video, audio, and PDF. Supported capabilities include code execution, computer use in preview, file search, function calling, Google Maps grounding, search grounding, structured outputs, thinking, and URL context. The model outputs text only and does not support image generation, audio generation, or the Live API. Google lists it as a previous-generation Flash model because Gemini 3.7 Flash and Gemini 3.8 Flash are newer models. Pricing is the standard comparison price published in July 2026; platform-specific pricing, caching rates, and inference-mode adjustments may differ. Editorial scores are comparative estimates rather than vendor specifications.

Cost

Model pricing

Input $1.50 per 1 million input tokens
Output $7.50 per 1 million output tokens
Model guide

Gemini 3.6 Flash: Google’s Fast, Long-Context Workhorse Model

Gemini 3.6 Flash is a stable, previous-generation Google model built for fast multimodal reasoning, coding, tool-using agents, and long-context applications. It accepts text, images, audio, video, and PDFs, supports up to 1,048,576 input tokens and 65,536 output tokens, and produces text output with function calling and structured-output support.

What is Gemini 3.6 Flash?

Gemini 3.6 Flash is a general-purpose multimodal model from Google. It is designed for applications that need fast responses while still handling substantial reasoning, coding, document analysis, and agentic tasks. In this context, “Flash” describes its positioning as a speed- and efficiency-oriented model rather than Google’s largest or most capability-focused option.

Google describes Gemini 3.6 Flash as a workhorse model for coding, knowledge work, multimodal understanding, spatial reasoning, and rapid agentic execution. The model is stable and currently accessible, although it is classified as a previous-generation Flash model because Gemini 3.7 Flash and Gemini 3.8 Flash are newer entries in the same broad lineup. No shutdown date has been announced.

The canonical stable model identifier is gemini-3.6-flash. It is available through the Gemini API, Google AI Studio, Google Cloud, and related Google distribution channels.

Supported inputs and outputs

Gemini 3.6 Flash accepts several types of input, which allows one request to combine ordinary instructions with files or media. Supported inputs include:

  • Text
  • Images
  • Audio
  • Video
  • PDF documents

Its native output is text. That includes normal written responses, code, tool calls, and supported structured responses. The model does not natively generate images, video, audio, speech, or music. This distinction matters when comparing it with a multimodal application that can both understand media and create media.

For example, Gemini 3.6 Flash can analyze a video, explain what happens in a PDF, inspect an image, or use information from an audio recording as part of a text response. It should not be selected when the application itself requires the model to return a generated image, video clip, spoken response, or soundtrack.

Context window and output limit

The model supports an input context window of up to 1,048,576 tokens, or approximately one million tokens. A context window is the amount of information the model can consider in a request, including instructions, conversation history, documents, code, and other supplied content. This large limit is useful for long source files, sizable codebases, extended transcripts, video-related analysis, and multi-step workflows where earlier information needs to remain available.

The maximum output is 65,536 tokens. This is a ceiling rather than a promise that every request will produce an output of that size. Actual responses depend on the prompt, generation settings, safety systems, tool activity, and the application’s own limits.

A large context window does not guarantee perfect attention to every detail. Long inputs can still contain ambiguous, repetitive, or conflicting information, and the model can still make factual or reasoning errors. The practical benefit is that developers can provide more relevant material in one interaction without dividing it into as many separate requests.

Reasoning, coding, and tool support

Gemini 3.6 Flash is intended for reasoning-heavy general-purpose work, including multi-step analysis and agentic execution. “Agentic” applications use a model to plan or carry out a sequence of actions, often by calling external tools and using the results in later steps. The model supports thinking, function calling, code execution, search grounding, URL context, file search, and Google Maps grounding. Computer use is supported in preview.

Function calling allows an application to define operations that the model can request, such as looking up an order, querying a database, or creating a task. The application remains responsible for executing those operations and controlling permissions. Search grounding and URL context can connect responses to external information during use, while file search helps the model work with an application's indexed material.

The model is also positioned for coding assistants and software workflows. Suitable tasks include explaining code, generating implementation drafts, reviewing files, tracing errors, transforming code, and helping coordinate multi-step development tasks. The supplied research describes coding as one of its stronger areas, but this is an editorial assessment rather than a vendor-published guarantee of accuracy.

Structured outputs are supported, allowing applications to request responses that follow a defined structure. This can make model responses easier to process programmatically. Structured output support should not be treated as proof of a separate, universal JSON mode: the exact behavior depends on the API feature and schema used by the application.

Where it fits in Google’s model lineup

Gemini 3.6 Flash sits between very large frontier-style models and simpler low-cost models in practical terms. Its design emphasizes a balance of response speed, multimodal understanding, long-context capacity, tool use, and cost. Google’s current catalog places it behind newer Gemini 3.7 Flash and Gemini 3.8 Flash models in generation order, but being previous-generation does not make it unsuitable for every workload.

Compared with a larger, slower model, Gemini 3.6 Flash may be a better fit when an application processes many requests, needs lower latency, or must repeatedly call tools. Compared with a smaller or less capable model, its broad media input support, one-million-token context, coding focus, and agent features may justify the additional cost and complexity.

The research gives Gemini 3.6 Flash editorial scores of 8 for reasoning, 8 for coding, 9 for speed, and 8 for cost. These are comparative estimates, not specifications published by Google. They summarize the model’s intended trade-off: strong general capability with particularly favorable speed characteristics, rather than maximum capability at any price.

Pricing and access

The supplied Google comparison information lists standard pricing of $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. Input tokens are the text or other processed content sent to the model; output tokens are the generated response. These prices should be treated as standard comparison pricing rather than a universal bill for every use case.

Actual charges can vary by API product, inference mode, caching status, region, and applicable platform terms. Batch and flex inference options are supported, as is priority inference, so developers should check the applicable Google documentation before estimating production costs. A request containing a very large context can also consume substantially more input tokens than a short prompt, even if the generated answer is brief.

Limitations and considerations

Gemini 3.6 Flash can hallucinate, meaning it may produce information that sounds plausible but is incorrect. It may also experience slowness or timeout issues. Tool support can improve access to current or application-specific information, but tools do not eliminate the need to validate model-generated plans, code, or actions.

The model-card knowledge cutoff is identified as March 2026. The supplied research notes that coverage can vary by domain and may be limited to January 2025 for some areas. Search grounding and other external tools can provide newer information during a request, but they do not change the underlying knowledge cutoff.

Gemini 3.6 Flash is not the right choice for native media generation. Applications that need generated images, video, audio, speech, or music require a different model or a separate generation service. It may also be less appropriate when the newest available Flash generation is required, when a dedicated realtime audio model is more suitable, or when the highest possible reasoning performance matters more than latency and cost.

When to choose Gemini 3.6 Flash

Gemini 3.6 Flash is a sensible choice when an application needs several of the following at once:

  • Fast responses for high-volume or interactive workloads
  • Text, image, audio, video, and PDF understanding
  • Long-context document, code, or media analysis
  • Coding assistance and software-development workflows
  • Function calling, code execution, search, or other tools
  • Structured responses for downstream application processing
  • Lower latency or lower cost than a larger frontier-oriented model

It is especially well suited to document-processing systems, coding assistants, research tools, enterprise workflows, multimodal classification or analysis, and agents that need to perform several tool calls quickly.

Choose a newer sibling model when access to the latest Gemini Flash generation is itself a requirement. Choose a different model or service when the application needs native image, video, audio, speech, or music generation. For safety-critical or highly consequential use, independent validation and application-level controls remain necessary because the model’s speed and tool access do not guarantee factual correctness or safe execution.

Bottom line

Gemini 3.6 Flash is a practical, stable Google model for fast multimodal understanding, coding, long-context processing, and tool-using applications. Its strongest technical differentiators are the one-million-token input context, support for multiple media types, 65,536-token output ceiling, and broad agent-oriented tool support. Its main compromises are that it is no longer Google’s newest Flash generation, it produces text rather than media, and it can still hallucinate or encounter operational issues. For applications that value speed and breadth over having the newest or most specialized model, it remains a capable option.


Answers to Frequently Asked Questions

How much does Gemini 3.6 Flash cost, and when should it be used?
The listed standard pricing is $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, although actual costs can vary by product, inference mode, caching, region, and platform terms. It is a good fit for fast, high-volume multimodal analysis, coding assistants, long-context document processing, and tool-using agents. A different model or service is preferable when native media generation, the newest Flash generation, realtime audio, or maximum reasoning performance is required.
What tools and capabilities does Gemini 3.6 Flash support?
Gemini 3.6 Flash supports thinking, function calling, code execution, search grounding, URL context, file search, and Google Maps grounding. Computer use is supported in preview. It also supports structured outputs for responses that follow an application-defined schema.
How large is Gemini 3.6 Flash’s context window?
Gemini 3.6 Flash supports an input context window of up to 1,048,576 tokens, or approximately one million tokens. Its maximum output is 65,536 tokens. The large context is useful for long documents, codebases, transcripts, and multi-step workflows, but it does not guarantee perfect accuracy or attention to every detail.
What is Gemini 3.6 Flash?
Gemini 3.6 Flash is Google’s general-purpose multimodal model designed for fast responses, coding, reasoning, document analysis, and agentic workflows. Its stable model identifier is gemini-3.6-flash, and it is available through the Gemini API, Google AI Studio, Google Cloud, and related Google channels.
What types of input and output does Gemini 3.6 Flash support?
Gemini 3.6 Flash accepts text, images, audio, video, and PDF documents. Its native output is text, including code, tool calls, and supported structured responses. It does not natively generate images, video, audio, speech, or music.


Sources 5
Provider

About Google DeepMind