Gemini 3

Gemini 3 Flash

by Google DeepMind · Preview

Gemini 3 Flash is Google DeepMind's preview reasoning model for fast, cost-conscious multimodal and agentic applications. It accepts text, images, video, audio, and PDFs; supports a 1-million-token context window and 65,536-token outputs; and provides function calling, search grounding, code execution, file search, computer use, structured outputs, and configurable thinking levels.

Text Reasoning Coding
Gemini 3 Flash is a preview model in Google's Gemini 3 family, released on December 17, 2025. It combines configurable reasoning with Flash-level speed and lower token pricing than Gemini 3 Pro. The model is designed for agentic workflows, everyday coding, planning, long-context document analysis, and applications that need to interpret several types of input while returning text.
Outputs

What Gemini 3 Flash can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3
Model type Multimodal
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff January 2025
Release date 2025-12-17
Status Preview
Knowledge cutoff notes

Google's Gemini 3 developer guide lists January 2025 as the knowledge cutoff for Gemini 3 Flash Preview. Search grounding and other external tools can provide newer information during use but do not change the underlying cutoff.

Model notes

Canonical API identifier is gemini-3-flash-preview. The model accepts text, images, video, audio, and PDF inputs and returns text. It is based on the Gemini 3 Pro reasoning foundation and supports configurable thinking levels. Provider documentation lists function calling, code execution, computer use, file search, Google Maps grounding, search grounding, structured outputs, URL context, caching, batch API, Flex inference, and Priority inference. Google reports a January 2025 knowledge cutoff. Pricing is listed per 1 million tokens and may change while the model remains in preview. The editorial scores are comparative estimates, not provider specifications.

Cost

Model pricing

Input $0.50 per 1 million input tokens
Output $3 per 1 million output tokens
Model guide

Gemini 3 Flash: Fast, Long-Context Reasoning for Agentic Workflows

Gemini 3 Flash is Google DeepMind's preview multimodal reasoning model for developers who need stronger planning, coding, and tool use than lightweight models provide, but lower latency and cost than a Pro-tier model. It accepts text, images, video, audio, and PDFs, supports a 1-million-token context window, produces up to 65,536 output tokens, and works with tools such as function calling, search grounding, code execution, file search, and computer use.

What is Gemini 3 Flash?

Gemini 3 Flash is a preview multimodal reasoning model from Google DeepMind. Its canonical API identifier is gemini-3-flash-preview. The model is based on the Gemini 3 Pro reasoning foundation, but is positioned for applications that need faster responses and lower cost than a Pro-tier model.

The model is not limited to simple short-answer chat. It is intended for agentic workflows, where an application gives the model access to tools and lets it plan several steps, call functions, inspect files, search for information, or execute code. It can also handle routine coding and reasoning tasks without requiring the higher cost and latency associated with a more capability-focused model.

Gemini 3 Flash remains a preview release. That status matters for production planning: Google may change its behavior, availability, rate limits, pricing, or endpoint lifecycle. Teams building important applications should monitor Google's current model documentation rather than assuming that preview characteristics will remain fixed.

Where Gemini 3 Flash fits in the Gemini lineup

Gemini 3 Flash occupies the speed-and-cost-oriented position in the Gemini 3 family. Compared with Gemini 3 Pro, its purpose is to provide a more economical and responsive option while retaining advanced reasoning, multimodal understanding, and tool support. This makes it a middle ground between basic lightweight models and a slower, more expensive model selected primarily for maximum capability.

The distinction is practical rather than merely a naming difference. A routine classification, coding suggestion, document question, or tool-routing task may not justify Pro-level cost. Conversely, a workflow that requires planning across a very large collection of files, interpreting images and audio, or coordinating multiple tools may need more than a minimal model can reliably provide. Gemini 3 Flash is designed for this intermediate workload.

Inputs, outputs, and context limits

Gemini 3 Flash accepts several input types:

  • Text
  • Images
  • Video
  • Audio
  • PDF files

Its published input context limit is up to 1,048,576 tokens, or approximately one million tokens. Context is the information the model can consider during a request, including the prompt, conversation history, and supplied documents. A million-token window is useful for large document collections, long code repositories, extended transcripts, and multimodal analysis that would otherwise require aggressive splitting.

The maximum output limit is 65,536 tokens. This is a ceiling rather than a promise that every request will produce a long response. In practice, applications should request only the amount of output they need, because larger responses can increase cost and latency.

Although Gemini 3 Flash accepts images, video, audio, and PDFs, its output is text. It does not natively generate images, audio, or video according to the supplied model specifications. Multimodal input should therefore not be confused with multimodal output.

Reasoning, planning, and coding

Gemini 3 Flash includes configurable thinking levels: minimal, low, medium, and high. These settings allow an application to choose a trade-off between response speed, token use, and depth of reasoning. Lower settings are better suited to routine requests where fast turnaround matters. Higher settings can be used for more complicated planning, analysis, or coding tasks, but may take longer and consume more tokens.

Google positions the model for reasoning and planning as well as everyday coding. It can be used to break down a task, propose an implementation, inspect a codebase, explain an error, or coordinate a sequence of tool calls. The large context window is especially relevant to repository analysis and long technical documents because the application can provide more surrounding information in one request.

Google reported a 78% SWE-bench Verified score for agentic coding in its Gemini CLI announcement and described Gemini 3 Flash as outperforming Gemini 2.5 Pro on several high-frequency development workflows. These are provider-reported results, not guarantees for every programming language, repository, prompt, or development environment. They should be treated as an indication of Google's intended positioning rather than a substitute for testing the model on a representative codebase.

Tools and agentic application support

Gemini 3 Flash supports function calling and tool use. Function calling allows the model to return a structured request for an application-defined operation, such as looking up an order, querying a database, or creating a calendar entry. The application—not the model—controls whether the operation is actually executed.

Provider documentation also lists support for search grounding, Google Maps grounding, code execution, file search, computer use, URL context, and structured outputs. Search grounding can help a request obtain newer external information than the model's internal knowledge. Code execution can support calculations or programmatic analysis, while file search can make relevant material available from a larger file collection. Computer use introduces additional operational risk because the surrounding system may allow the model to interact with software interfaces.

These features should be permissioned carefully. A model-generated tool call should be validated, logged, and restricted according to the consequences of the action. Reading a document is a lower-risk operation than sending a message, changing a record, executing code with broad permissions, or controlling a computer. Gemini 3 Flash's tool support expands what an application can do, but it does not remove the need for application-level security and approval rules.

Structured outputs can help applications receive data in a defined format instead of extracting fields from free-form prose. This is useful for tasks such as returning an invoice record, a list of code issues, or a workflow plan. Structured output support should not automatically be interpreted as a separate general-purpose JSON mode; the supplied research records the model's JSON-mode field as unknown while separately listing structured outputs as supported.

Pricing and consumption

The published Gemini 3 pricing table lists Gemini 3 Flash Preview at $0.50 per 1 million input tokens and $3 per 1 million output tokens. Input tokens are the text and other request content sent to the model, while output tokens are the generated response. Applications that produce lengthy responses or use high thinking levels should account for output consumption as well as input volume.

The model is available through Google AI Studio and the Gemini API, with preview access also available through Google Cloud and related Google developer products. The supplied documentation lists support for batch API, Flex inference, and Priority inference options. The appropriate consumption mode depends on whether an application values lower-cost asynchronous processing, flexible capacity, or more predictable performance. Availability and terms can change while the model remains in preview.

Pricing alone does not determine total operating cost. Tool calls, search or other external services, repeated context, retries, and application-side processing can contribute to the overall cost of a workflow. Caching support may help reduce repeated-context overhead in suitable applications, but developers should confirm the current pricing and caching rules before making detailed cost projections.

Main strengths and limitations

Strengths

  • Large context: The one-million-token input limit is well suited to long documents, repositories, transcripts, and mixed-media investigations.
  • Broad input handling: Text, image, video, audio, and PDF support allows one model to analyze different forms of source material.
  • Configurable reasoning: Thinking levels let applications prioritize speed for simple requests or deeper processing for complex tasks.
  • Agentic support: Function calling, search grounding, code execution, file search, and computer use enable workflows that go beyond text generation.
  • Lower-cost positioning: Its published token prices are intended to make advanced reasoning more economical than a Pro-tier model.
  • Coding focus: Google specifically positions the model for everyday coding and agentic development workflows.

Limitations

  • Preview status: The model's interface, behavior, pricing, limits, and availability may change.
  • Text-only output: It does not directly generate images, audio, or video.
  • Latency and cost vary with reasoning: Higher thinking levels can increase processing time and token consumption.
  • Knowledge cutoff: Google's developer guide lists January 2025 as the underlying knowledge cutoff. Search grounding may supply newer information during a request, but it does not change that cutoff.
  • Tool risk: Search, code execution, computer use, and external actions require careful permissions and validation.
  • Benchmark limits: Provider-reported coding results may not predict performance on an individual team's codebase or workflow.

When to choose Gemini 3 Flash

Choose Gemini 3 Flash when an application needs a combination of multimodal understanding, meaningful reasoning, tool use, and fast or cost-conscious operation. Good candidates include coding assistants, repository and document analysis, research workflows with search grounding, structured extraction from PDFs, planning systems, customer-support triage, and agents that need to inspect several types of input before taking a controlled action.

It is particularly suitable when the model must process a large amount of context but does not need image, audio, or video generation. For example, a development assistant could review a large repository, explain a failing test, call a search or file tool, and return a proposed patch. A research workflow could combine PDFs, images, and web-grounded information, then produce a structured report.

Another option may be more appropriate when the application requires native media generation, speech-to-speech interaction, or a stable non-preview contract. A higher-tier or more capability-focused model may be preferable for tasks where maximum reasoning quality is more important than speed and price. Conversely, a simpler lightweight model may be a better fit for short, repetitive, low-risk requests where multimodal analysis and advanced tool use are unnecessary.

Bottom line

Gemini 3 Flash is a fast, cost-oriented member of the Gemini 3 family that retains substantial reasoning and agentic functionality. Its defining practical combination is a one-million-token context window, broad multimodal input, configurable thinking, tool support, and lower published pricing than a Pro-tier model. The main compromises are preview status, text-only output, variable reasoning cost and latency, and the need to secure any connected tools.


Answers to Frequently Asked Questions

When should you choose Gemini 3 Flash instead of another Gemini model?
Choose Gemini 3 Flash when you need multimodal input, substantial reasoning, tool use, large-context analysis, and relatively fast or cost-conscious operation. It is suitable for coding assistants, document and repository analysis, research workflows, structured PDF extraction, and controlled agents. A different model may be better for native media generation, speech-to-speech interaction, maximum reasoning quality, or a stable non-preview contract.
What tools and agentic capabilities does Gemini 3 Flash support?
Gemini 3 Flash supports function calling, search grounding, Google Maps grounding, code execution, file search, computer use, URL context, and structured outputs. Tool calls should be validated, logged, permissioned, and restricted by application-level security controls before execution.
How much does Gemini 3 Flash cost?
The published Gemini 3 pricing table lists Gemini 3 Flash Preview at $0.50 per 1 million input tokens and $3 per 1 million output tokens. Actual workflow costs can also include tool calls, external services, repeated context, retries, and application-side processing.
What is Gemini 3 Flash and what is its API identifier?
Gemini 3 Flash is a preview multimodal reasoning model from Google DeepMind designed for fast, cost-conscious applications and agentic workflows. Its canonical API identifier is gemini-3-flash-preview.
What are Gemini 3 Flash's context and output limits?
Gemini 3 Flash supports up to 1,048,576 input tokens, or approximately one million tokens, and a maximum output of 65,536 tokens. It accepts text, images, video, audio, and PDF files, but its native output is text only.


Sources 5
Provider

About Google DeepMind