Gemini 3.5

Gemini 3.5 Flash

by Google DeepMind · Generally available; stable

Google DeepMind’s Gemini 3.5 Flash is a stable model for agentic workflows, coding agents, long-context analysis, structured outputs, and tool use. It accepts text, image, video, audio, and PDF inputs, supports up to 1,048,576 input tokens and 65,536 output tokens, and offers standard, cached, batch, Flex, and Priority pricing modes. Its main limitations are text-only output, no Live API support, preview computer use, and a January 2025 knowledge cutoff.

Text Actions Reasoning Coding
Gemini 3.5 Flash is designed for developers who need a fast, capable model for agents, coding systems, document analysis, and other workflows that may require many steps or large amounts of context. It accepts text, images, video, audio, and PDF inputs, but produces text rather than images, audio, or video. Its main trade-off is a higher standard price than simpler Flash models, balanced by stronger reasoning, coding, and tool-use capabilities.
Outputs

What Gemini 3.5 Flash can produce

Text Actions
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.5
Model type General Purpose
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff January 2025
Release date 2026-05-19
Status Generally available; stable
Knowledge cutoff notes

Google's Gemini 3.5 Flash documentation states that the model has a knowledge cutoff of January 2025. Search grounding can provide newer information during use but does not change the underlying cutoff.

Model notes

Canonical model ID is gemini-3.5-flash. The model is generally available and stable as of September 26, 2026. It supports thinking, structured outputs, function calling, code execution, file search, URL context, Google Search grounding, Google Maps grounding, and computer use in preview. Standard context caching is priced at $0.15 per 1M cached tokens, with separate storage charges. Standard output pricing includes thinking tokens. The model is behind the gemini-flash-latest alias, which is not a separate model entity. Editorial scores are comparative estimates rather than vendor-published ratings.

Cost

Model pricing

Input $1.50 per 1M tokens standard; $0.75 per 1M tokens Batch and Flex; $2.70 per 1M tokens Priority
Output $9.00 per 1M tokens standard; $4.50 per 1M tokens Batch and Flex; $16.20 per 1M tokens Priority
Model guide

Gemini 3.5 Flash: Google’s Stable Model for Agentic Coding and Long-Context Work

Gemini 3.5 Flash is Google DeepMind’s generally available, stable Flash model for high-throughput agentic workflows, iterative coding, long-context analysis, and multimodal understanding. It combines a 1,048,576-token input context window, up to 65,536 output tokens, configurable thinking, tool use, structured outputs, and production features such as caching and batch processing.

What is Gemini 3.5 Flash?

Gemini 3.5 Flash is a generally available model from Google DeepMind. Its stable model identifier is gemini-3.5-flash. The model sits in Google’s Gemini 3.5 family and is positioned as a high-capability Flash model: it is intended to deliver stronger reasoning and coding than a basic low-cost model while retaining the responsiveness and deployment flexibility associated with the Flash line.

The model is primarily aimed at applications that must perform several related tasks rather than answer a single short question. Examples include an agent that plans a task, calls tools, checks the results, and continues; a coding assistant that repeatedly edits and tests a repository; or a document-analysis system that compares large collections of files.

Google lists Gemini 3.5 Flash as stable and generally available. The release date supplied for the model is May 19, 2026. The model’s documented knowledge cutoff is January 2025, so information about later events requires an external grounding source such as Google Search.

Key specifications at a glance

SpecificationGemini 3.5 Flash
ProviderGoogle DeepMind
Model IDgemini-3.5-flash
StatusGenerally available and stable
Input context1,048,576 tokens
Maximum output65,536 tokens
Knowledge cutoffJanuary 2025
Input modalitiesText, images, video, audio, and PDF
Output modalityText, including structured text when configured
ThinkingConfigurable

A token is a small unit of text used by language models. The one-million-token input limit is large enough for substantial source material, although the usable amount depends on the format of the content and the complexity of the requested task. The 65,536-token output limit includes thinking tokens where applicable, according to the supplied model documentation.

Reasoning, coding, and agent work

Gemini 3.5 Flash supports configurable thinking. In practical terms, an application can use the model’s reasoning capacity for tasks that benefit from planning, intermediate checking, or multi-step problem solving, while avoiding unnecessary processing for simpler requests. The research describes the model as particularly suited to long-horizon work, where a useful result depends on a sequence of decisions rather than one generation.

Its coding role includes iterative software development, repository analysis, code generation, and tool-assisted workflows. A coding agent could use the model to inspect files, propose a change, call a testing or execution tool, interpret the result, and revise its output. The model supports code execution, function calling, file search, URL context, and other mechanisms that allow an application to connect model responses with external actions or information.

The supplied comparative assessment gives Gemini 3.5 Flash a reasoning score of 9 and a coding score of 9, but these are editorial estimates rather than provider-published benchmark results. They should be treated as directional evaluations, not guaranteed performance measurements for a particular workload.

Multimodal input, text output

Gemini 3.5 Flash can accept several kinds of input: ordinary text, images, video, audio, and PDF documents. This makes it suitable for tasks such as extracting information from a scanned document, reviewing a video transcript or scene sequence, analyzing spoken content, or combining written instructions with visual evidence.

Its output is text-only. That text can be ordinary prose, code, or structured machine-readable data when structured outputs are configured. Gemini 3.5 Flash does not natively generate images, audio, or video. This distinction matters when selecting a model: multimodal input does not mean that the model can produce every media type.

The model also does not support the Live API according to the supplied research. Applications needing native realtime voice interaction or media generation should therefore consider a different model or service rather than treating Gemini 3.5 Flash as a general-purpose media-generation system.

Tools and production features

Gemini 3.5 Flash supports structured outputs, function calling, code execution, file search, URL context, Google Search grounding, and Google Maps grounding. Computer use is listed as a preview capability. These features allow developers to build systems that can retrieve information, interact with services, execute code, or carry out controlled interface actions instead of relying only on the model’s internal knowledge.

Structured outputs are useful when the response must conform to a defined machine-readable shape, such as a list of extracted fields or a workflow decision. Function calling allows the application to expose operations that the model may request, while the application remains responsible for executing and validating those operations. These capabilities are especially relevant to agents, but they also introduce the need for permission controls, input validation, monitoring, and careful handling of failures.

For larger deployments, the model supports context caching, batch processing, Flex inference, and Priority inference. Caching can reduce repeated input costs when the same context is reused, while batch processing is intended for workloads that do not require immediate responses. Priority inference provides a faster service tier at a higher price. Computer use is still preview functionality, so production teams should evaluate its reliability and safety separately before depending on it for important operations.

Pricing and speed-cost trade-offs

Standard paid pricing is $1.50 per 1 million input tokens and $9.00 per 1 million output tokens. The output rate includes thinking tokens. Context caching is priced at $0.15 per 1 million cached tokens, with separate storage charges.

Batch and Flex pricing reduce the standard rates to $0.75 per 1 million input tokens and $4.50 per 1 million output tokens. Priority inference is listed at $2.70 per 1 million input tokens and $16.20 per 1 million output tokens. These are different service modes rather than different model identities, so the appropriate choice depends on whether the application prioritizes cost reduction, ordinary production access, or faster handling.

The cost calculation should include both the information sent to the model and the information it produces. Long prompts, large files, repeated context, and extensive reasoning can all affect usage. For recurring analysis of the same large instructions or reference material, caching may be useful. For offline classification, document processing, or other jobs that can wait, Batch or Flex pricing may be more economical than standard or Priority inference.

Best use cases

  • Agentic workflows: Multi-step assistants that plan tasks, call functions, inspect results, and continue until a goal is reached.
  • Coding agents: Repository exploration, iterative code changes, debugging, test interpretation, and software-development assistance.
  • Long-context analysis: Large documents, PDFs, video, audio, and mixed-media collections that require cross-referencing.
  • Tool-using enterprise workflows: Applications that combine model reasoning with search, maps, file retrieval, code execution, or business functions.
  • Scaled production systems: Applications that need stronger reasoning than a basic fast model and can use caching, batch processing, Flex inference, or Priority inference.

The one-million-token context window is particularly relevant when an application must preserve a large working set. However, a large context limit is not a guarantee that every task will be handled perfectly. Developers should still retrieve the most relevant material, manage prompts carefully, and test how the model behaves when information is spread across a very large input.

Limitations and when to choose another option

Gemini 3.5 Flash is not the right choice for every application. It cannot natively generate images, audio, or video, so media-generation workloads require another option. It also does not support the Live API, making it unsuitable for applications built around that specific realtime interface.

The January 2025 knowledge cutoff is another important limitation. The model can use Google Search grounding and other external tools for newer information, but its built-in knowledge does not automatically become current. Applications that need reliable, up-to-date facts should explicitly design for retrieval, grounding, source checking, and failure handling.

Its standard output price is substantially higher than its input price, and extensive reasoning or long responses can increase the bill. A smaller or simpler fast model may be more appropriate for short classification, routine extraction, simple rewriting, or high-volume requests where advanced reasoning is unnecessary. Conversely, a specialized media model is more suitable when the required output is an image, sound, or video rather than text.

Computer use is available only in preview, which makes it less appropriate for workflows that require a mature, highly predictable interface-control capability. In all tool-enabled deployments, the application should restrict permissions and verify model-generated actions before they affect external systems.

When to choose Gemini 3.5 Flash

Choose Gemini 3.5 Flash when the main challenge is combining strong reasoning, coding, multimodal understanding, long context, and tool use in one text-producing model. It is a good fit when a workflow may need to inspect substantial material, make a plan, call external functions, and revise its answer over multiple steps.

Choose a lower-cost or simpler model when response speed and price matter more than advanced reasoning, especially for predictable, short, repetitive tasks. Choose a model with native image, audio, or video output when generation of those media types is central to the product. Choose an option with current information built in, or add a reliable grounding layer, when answers depend on events after January 2025.

Overall, Gemini 3.5 Flash occupies a middle ground between basic high-throughput models and more specialized or media-oriented systems. Its value comes from the combination of a very large input context, configurable thinking, strong coding and agent support, and multiple production inference modes. Those benefits are most worthwhile when the application genuinely uses them; they may not justify the cost for straightforward text processing.


Answers to Frequently Asked Questions

How much does Gemini 3.5 Flash cost?
Standard paid pricing is $1.50 per 1 million input tokens and $9.00 per 1 million output tokens. Batch and Flex pricing is $0.75 per 1 million input tokens and $4.50 per 1 million output tokens, while Priority inference costs $2.70 per 1 million input tokens and $16.20 per 1 million output tokens. Cached tokens cost $0.15 per 1 million tokens, excluding separate storage charges.
Does Gemini 3.5 Flash generate images, audio, or video?
No. Gemini 3.5 Flash accepts text, images, video, audio, and PDF input, but its output is text-only, including code or structured text. It does not natively generate images, audio, or video and does not support the Live API.
What can Gemini 3.5 Flash be used for?
It is well suited to multi-step agents, coding assistants, repository analysis, debugging, large-document processing, PDF and media analysis, and enterprise workflows that use function calling, code execution, file search, URL context, Google Search grounding, or Google Maps grounding.
What is Gemini 3.5 Flash?
Gemini 3.5 Flash is a generally available and stable Google DeepMind model identified as gemini-3.5-flash. It is designed for agentic workflows, coding, long-context analysis, multimodal understanding, and tool-assisted applications.
What are the context window and output limits of Gemini 3.5 Flash?
Gemini 3.5 Flash supports up to 1,048,576 input tokens and a maximum output of 65,536 tokens. The output limit includes thinking tokens where applicable.


Sources 5
Provider

About Google DeepMind