Gemini 3

Gemini 3.7 Flash

by Google DeepMind · Generally available; previous-generation Flash model, currently supported

Gemini 3.7 Flash is a generally available Google DeepMind model for coding, tool-using agents, long-context analysis and enterprise automation. It accepts text, images, video, audio and PDFs, supports a 1,048,576-token context window, configurable thinking, structured outputs, search grounding, caching and batch processing, and produces text output at introductory prices designed for cost-conscious, high-throughput applications.

Text Reasoning Coding
Gemini 3.7 Flash is a high-speed model in Google’s Gemini 3 family, released generally on August 13, 2026. It is designed for software engineering, agentic workflows, structured extraction and multimodal reasoning rather than media generation. The model combines a 1-million-token context window with function calling, search grounding, code execution, caching and batch processing, while keeping introductory API prices below those of larger frontier models.
Outputs

What Gemini 3.7 Flash can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3
Model type General Purpose
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff 2026-03
Release date 2026-08-13
Status Generally available; previous-generation Flash model, currently supported
Knowledge cutoff notes

Google DeepMind’s model card states that the knowledge cutoff is March 2026. Search grounding and other retrieval tools may provide newer information during use but do not alter the underlying cutoff.

Model notes

Canonical model ID is gemini-3.7-flash. Google describes it as a previous-generation Flash model after Gemini 3.8 Flash but lists it as fully supported. Thinking supports low, medium and high settings; minimal thinking is not supported. Computer use is supported in preview. Audio generation, image generation and Live API are not supported. Standard introductory pricing applies through December 31, 2026; batch and Flex inference are listed at 50% of standard rates. Search grounding, Google Maps grounding, URL context, code execution, file search, function calling and structured outputs are supported. The documented knowledge cutoff is March 2026.

Cost

Model pricing

Input $0.75 per 1M tokens through 2026-12-31; $1.50 per 1M tokens starting 2027-01-01
Output $3.75 per 1M tokens, including thinking tokens, through 2026-12-31; $7.50 per 1M tokens starting 2027-01-01
Model guide

Gemini 3.7 Flash: Google’s Fast, Cost-Efficient Model for Coding and AI Agents

Gemini 3.7 Flash is Google DeepMind’s generally available, natively multimodal model for coding, tool-using agents, long-context analysis and enterprise automation. It accepts text, images, video, audio and PDFs, provides a 1,048,576-token context window, supports configurable thinking and produces text output. Its combination of relatively low introductory pricing, high throughput and developer tools makes it a practical choice for applications that need more capability than a lightweight model but do not require a larger, more expensive frontier model.

What is Gemini 3.7 Flash?

Gemini 3.7 Flash is a generally available model from Google DeepMind. The “Flash” positioning indicates a focus on speed, efficiency and cost, while the model remains capable of handling complex, multi-step work. Google describes it as a model for coding, agentic workflows, tool use and reliable execution of tasks that require several stages of reasoning.

In practical terms, Gemini 3.7 Flash is intended to sit between lightweight models optimized mainly for simple, inexpensive requests and larger frontier models aimed at maximum capability. It is especially relevant when an application must process substantial context or call external tools repeatedly, but still needs predictable throughput and controlled inference costs.

The stable model identifier is gemini-3.7-flash. It is a previous-generation Flash model relative to Gemini 3.8 Flash, but the supplied documentation lists Gemini 3.7 Flash as fully supported and generally available.

Inputs, outputs and context limits

Gemini 3.7 Flash accepts several types of input:

  • Text
  • Images
  • Video
  • Audio
  • PDF documents

This makes it suitable for applications that need to combine written instructions with visual, spoken or document-based information. Examples include reviewing a software design document with diagrams, extracting information from long PDFs, examining video content, or combining an audio recording with a written task.

The model supports up to 1,048,576 input tokens in its context window. A token is a unit of text or other encoded content; the exact number of words represented by a token varies by language and content type. A context window of this size allows developers to provide very large documents, codebases or multimodal inputs in one request, subject to the relevant API handling and content limits.

The maximum output length is 65,536 tokens. Gemini 3.7 Flash produces text output. It does not natively generate images, video or audio, so its multimodal capability is primarily about understanding and reasoning over different input formats rather than directly creating media.

Reasoning and coding capabilities

Gemini 3.7 Flash supports configurable thinking at low, medium and high levels. Thinking allows the model to spend additional internal processing on a task before producing its answer. The minimal thinking setting is not supported, so developers should account for the available low, medium and high choices when tuning latency, cost and answer quality.

The model is aimed particularly at software engineering. It can generate and explain code, analyze repositories or files, help diagnose implementation problems, and participate in workflows where code is created, tested or revised through tools. The supplied research rates its coding capability and speed highly as editorial evaluations, not as scores published by Google. Those assessments reflect the model’s intended use and documented feature set rather than a specific benchmark result.

For structured extraction, Gemini 3.7 Flash supports structured outputs. This can help an application request machine-readable responses that follow a defined schema instead of relying on free-form prose. Structured output is useful for tasks such as extracting fields from invoices, converting documents into records, classifying support requests or returning consistent action parameters. It should not be treated as a guarantee that every response is factually correct or that application-side validation is unnecessary.

Tools and agent workflows

Gemini 3.7 Flash supports function calling, which allows a developer to describe external functions and let the model request their use. The application, rather than the model itself, normally executes the function and returns the result. This pattern is useful for agents that need to query databases, call business systems, retrieve records or perform controlled actions.

Documented developer capabilities include:

  • Function calling and tool use
  • Code execution
  • File search
  • URL context
  • Google Search grounding
  • Google Maps grounding
  • Structured outputs
  • Implicit context caching
  • Batch processing
  • Flex inference
  • Priority inference

Computer use is available as a preview capability. That status matters for production planning: preview features can have changing behavior, availability or operational constraints, so teams should avoid assuming that computer control will remain identical over time.

Search grounding and other retrieval tools can provide information newer than the model’s internal knowledge. The documented knowledge cutoff is March 2026, but retrieval during a request does not change that underlying cutoff. Developers should distinguish between what the model already knows and facts supplied by a search or other external tool.

Pricing and cost trade-offs

As of September 26, 2026, standard paid Gemini API pricing is $0.75 per 1 million input tokens and $3.75 per 1 million output tokens, including thinking tokens. Google lists these as introductory prices through December 31, 2026. From January 1, 2027, the listed standard prices increase to $1.50 per 1 million input tokens and $7.50 per 1 million output tokens.

Batch and Flex inference are listed at 50% of the standard introductory rates through December 31, 2026: $0.375 per 1 million input tokens and $1.875 per 1 million output tokens. These lower rates may be relevant for workloads that can tolerate deferred processing or use the applicable Flex service rather than requiring the most immediate response path.

Context caching is supported, with separate cached-input and cache-storage charges listed by Google. Caching can be useful when an application repeatedly sends a large, mostly unchanged instruction set, document collection or codebase. The financial benefit depends on the request pattern and the provider’s current cache pricing, so developers should calculate costs using the actual input, output, storage and cache-hit behavior of their application.

Search grounding is available for paid usage. Google applies a shared Gemini 3 search-request allowance and charges for applicable requests after that allowance is used. Search-related costs should therefore be considered separately from ordinary input and output token charges.

Speed versus capability

Gemini 3.7 Flash’s main practical trade-off is its balance of speed, context capacity and cost. It is designed for high-throughput applications that need more than basic text completion, including systems that analyze long inputs, call tools or perform several reasoning steps. The introductory token prices are also positioned below what a larger, more capability-focused model might cost.

That balance does not mean it is the best choice for every task. If an application prioritizes maximum reasoning performance over latency and cost, a larger frontier model may be more appropriate. Conversely, if a task is a short classification or simple transformation, a smaller model could be more economical. The supplied research does not provide a direct benchmark comparison with Gemini 3.8 Flash or other models, so decisions should be validated against the application’s own prompts, latency targets and error tolerance.

Best use cases

Gemini 3.7 Flash is a strong candidate for applications that combine substantial context with repeated or tool-assisted operations. Suitable use cases include:

  • Software engineering: code generation, repository analysis, debugging assistance and documentation work.
  • AI agents: workflows that call functions, retrieve information and complete multi-step tasks.
  • Long-context analysis: reviewing large documents, code collections, PDFs, videos or mixed-media materials.
  • Structured extraction: converting unstructured documents or conversations into schema-based records.
  • Enterprise automation: orchestrating business workflows that require search, file access, code execution or external systems.
  • High-volume processing: applications where Flash-level speed and introductory pricing matter more than maximum frontier-model capability.

For example, an enterprise application could provide Gemini 3.7 Flash with a long policy document, ask it to locate relevant clauses, call a search or internal retrieval function, and return the result in a fixed schema. A development tool could supply a large code context, ask for a change plan, and use function calling or code execution as part of an iterative workflow.

Limitations to consider

The most important limitation is that Gemini 3.7 Flash is an understanding and text-generation model, not a native media-generation model. It cannot directly generate images, video or audio, and it does not provide audio-to-audio interaction through the Live API. Applications requiring speech output, image creation or video generation should use a model or service designed for that output type.

Computer use is preview-only. It may be useful for experimentation and selected workflows, but teams should treat it differently from a mature, stable production contract. The model can also hallucinate, meaning it may produce plausible but incorrect information. Search grounding and application-side validation can reduce some risks but do not eliminate them.

Google’s documentation also notes that occasional slowness or timeout issues can occur. Long contexts, extended thinking, tool calls and large outputs may affect response behavior. Production systems should use timeouts, retries where appropriate, validation, permissions boundaries and clear handling for incomplete or failed tool calls.

Finally, the introductory prices are time-limited. The listed rates increase on January 1, 2027, so cost projections should not assume that the December 2026 pricing remains permanent.

When to choose Gemini 3.7 Flash

Choose Gemini 3.7 Flash when you need a generally available model that can understand text and multiple media types, work with a very large context, use external tools and produce structured text at relatively low introductory token prices. It is particularly well suited to coding systems, enterprise agents, document-heavy workflows and high-throughput multimodal analysis.

Consider another option when the application must generate images, video or audio; requires low-latency audio-to-audio conversation; depends on a stable non-preview computer-use feature; or prioritizes the highest possible reasoning capability regardless of cost and speed. A smaller model may be preferable for simple, repetitive requests, while a larger model may be preferable for the most demanding reasoning tasks. Gemini 3.7 Flash is most compelling when its combination of long context, tool support, coding focus and Flash-oriented efficiency matches the workload.


Answers to Frequently Asked Questions

What limitations should developers consider when using Gemini 3.7 Flash?
Gemini 3.7 Flash cannot directly generate images, video or audio and does not support audio-to-audio interaction through the Live API. Computer use is a preview feature, and the model can still produce inaccurate information. Long contexts, extended thinking and tool calls may also cause slowness or timeouts, so production applications should use validation, timeouts, retries and permission controls.
What are the main use cases for Gemini 3.7 Flash?
Common use cases include code generation, repository analysis, debugging, long-document and multimodal analysis, structured data extraction, enterprise automation, and AI agents that use functions, search, file access or other external tools.
How much does Gemini 3.7 Flash cost?
As of September 26, 2026, standard paid pricing is $0.75 per 1 million input tokens and $3.75 per 1 million output tokens, including thinking tokens. These introductory prices are listed through December 31, 2026, after which the listed rates increase to $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. Batch and Flex inference are listed at 50% of the introductory standard rates through December 31, 2026.
What is Gemini 3.7 Flash?
Gemini 3.7 Flash is a generally available Google DeepMind model designed for speed, efficiency, coding, AI agents, tool use and multi-step workflows. Its stable model identifier is `gemini-3.7-flash`.
What types of input does Gemini 3.7 Flash support?
Gemini 3.7 Flash accepts text, images, video, audio and PDF documents. It supports up to 1,048,576 input tokens and can generate up to 65,536 tokens of text output, but it does not natively generate images, video or audio.


Sources 6
Provider

About Google DeepMind