Gemini 3

Gemini 3.8 Flash

by Google DeepMind · Generally available

Google DeepMind’s Gemini 3.8 Flash is a generally available model for long-horizon software engineering, autonomous agents, multimodal document analysis, and enterprise workflows. It accepts text, images, video, audio, and PDFs, supports a 1,048,576-token context window and 65,536-token outputs, and offers configurable thinking, function calling, code execution, search grounding, file search, caching, and batch processing. It produces text only, so it is not intended for native image, audio, video, or speech generation.

Text Reasoning Coding
Gemini 3.8 Flash is a generally available multimodal model from Google DeepMind, released on September 2, 2026. It accepts text, images, video, audio, and PDF inputs but produces text output. Its main role is to handle demanding agentic and software-engineering tasks with a balance of reasoning ability, speed, and operating cost.
Outputs

What Gemini 3.8 Flash can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3
Model type Multimodal
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff March 2026
Release date September 2, 2026
Status Generally available
Knowledge cutoff notes

Google’s model card states a March 2026 knowledge cutoff and cautions that information in some domains may remain limited to January 2025, consistent with the Gemini 3 model family. Search grounding can supply newer external information during use but does not change the underlying cutoff.

Model notes

Canonical model ID is gemini-3.8-flash. The model is generally available and supports low, medium, and high thinking effort; minimal thinking is not supported. Computer use is supported in preview. Google documents text, image, video, audio, and PDF inputs with text output only. Standard introductory pricing applies through December 31, 2026, with scheduled standard pricing from January 1, 2027. Context caching, batch, Flex, and Priority inference are supported. The model card gives a March 2026 knowledge cutoff but notes that some domains may remain limited to January 2025. JSON mode is left unknown because structured outputs are documented separately and do not by themselves establish legacy JSON-mode support.

Cost

Model pricing

Input $0.75 per 1 million tokens through December 31, 2026; $1.50 per 1 million tokens from January 1, 2027. Batch/Flex introductory input price: $0.375 per 1 million tokens.
Output $3.75 per 1 million tokens, including thinking tokens, through December 31, 2026; $7.50 per 1 million tokens from January 1, 2027. Batch/Flex introductory output price: $1.875 per 1 million tokens.
Model guide

Gemini 3.8 Flash: Google’s Long-Context Model for Agents and Software Engineering

Gemini 3.8 Flash is Google DeepMind’s generally available Flash model for long-horizon software engineering, autonomous agents, multimodal understanding, and complex enterprise workflows. It combines a 1,048,576-token input context window, 65,536-token maximum output, configurable reasoning effort, tool use, search grounding, code execution, caching, and batch processing while returning text only.

What is Gemini 3.8 Flash?

Gemini 3.8 Flash is Google DeepMind’s generally available Flash model for long-running software engineering, autonomous agents, multimodal analysis, and complex knowledge workflows. It is designed for situations where a lightweight model may not provide enough reasoning or tool-use capability, but where a larger and more expensive frontier model would be unnecessary.

The model became generally available on September 2, 2026. It can be accessed through Google’s Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, Gemini applications, and other Google distribution channels. The canonical model identifier is gemini-3.8-flash.

“Flash” describes the model’s intended position: it emphasizes practical response speed and lower cost while retaining capabilities for multi-step reasoning, coding, and external tool use. This is a positioning distinction rather than a guarantee that every request will be fast. Google notes that occasional slowness or timeouts can occur, and higher reasoning settings can increase token usage and processing time.

Who should use Gemini 3.8 Flash?

Gemini 3.8 Flash is most suitable for developers and organizations building applications that need to process substantial context, call tools, inspect code, or complete multi-step tasks. Examples include repository analysis, long-running software-engineering agents, enterprise research assistants, document workflows, and applications that combine search with code execution or structured data processing.

It can also be useful for multimodal knowledge work. A request may combine text with images, video, audio, or PDF files, allowing the model to analyze different kinds of material in one workflow. The output remains text, so applications that require generated images, audio, video, or speech need a different model or an additional generation system.

Modalities and core capabilities

Verified documentation describes Gemini 3.8 Flash as accepting the following input types:

  • Text
  • Images
  • Video
  • Audio
  • PDF files

Its native output type is text. Although the model is multimodal on the input side, it is not a direct image, audio, video, music, or speech-generation model. This distinction matters when selecting it for an application: it can describe or reason about supplied media, but it does not produce those media formats as its native response.

The model supports configurable thinking effort at low, medium, and high levels. The minimal thinking setting is not supported. Thinking tokens are included in the documented output pricing, and selecting a higher effort can result in more extensive reasoning and greater token consumption. The research does not establish that a particular setting will always be superior, so developers should evaluate the trade-off on their own workload.

Gemini 3.8 Flash also supports function calling, code execution, file search, URL context, Google Search grounding, Google Maps grounding, structured outputs, context caching, batch processing, Flex inference, and Priority inference. Computer use is listed as a preview capability. These features depend on the surrounding API configuration and the tools enabled by the application; their presence does not mean the model independently performs every external action without developer integration.

Context window and output limit

The model has a documented input context limit of 1,048,576 tokens, or roughly one million tokens. A context window is the amount of information the model can consider in a request and its surrounding conversation. This large limit is intended for tasks such as examining substantial code repositories, processing long documents, combining multiple research files, or maintaining the working context of a multi-step agent.

The maximum output capacity is 65,536 tokens. This is a ceiling rather than a requirement: most requests will produce much shorter answers, and the actual response length depends on the prompt, configured thinking effort, tool activity, and application limits.

A large context window does not guarantee perfect recall or equally strong reasoning over every part of a very long input. Important instructions and evidence should still be organized clearly, and applications should test how the model behaves with their own document, code, and retrieval patterns.

Reasoning, coding, and agent work

Google positions Gemini 3.8 Flash for long-horizon software engineering and autonomous agents. In practical terms, this means the model is intended for workflows that require several connected steps rather than a single short answer. It can reason over a large working context, decide when tools are needed, call configured functions, use search or file retrieval, and incorporate returned information into a later response.

For coding tasks, the model can analyze repositories, generate or modify code, use code execution where enabled, and support software-engineering workflows that involve more than producing an isolated code snippet. Its large context is particularly relevant when an agent must inspect multiple files or preserve details across a longer task.

Tool use does not remove the need for application safeguards. Functions should validate arguments, code execution should run in an appropriately restricted environment, and search-grounded results should be checked before they are used in consequential workflows. The model can still produce incorrect reasoning or hallucinated information, even when tools are available.

Google’s published evaluations report improvements over Gemini 3.7 Flash on several agentic and reasoning benchmarks. Those results are benchmark-specific and should be treated as provider-reported evaluation evidence rather than a universal performance guarantee. The supplied research does not provide enough detail to claim that Gemini 3.8 Flash is superior on every task or against every competing model.

Pricing and processing options

For the Gemini API paid standard tier, introductory pricing through December 31, 2026 is $0.75 per 1 million input tokens and $3.75 per 1 million output tokens, including thinking tokens. Standard pricing is scheduled to rise on January 1, 2027 to $1.50 per 1 million input tokens and $7.50 per 1 million output tokens.

Batch and Flex processing are documented at half of the introductory standard rates through December 31, 2026: $0.375 per 1 million input tokens and $1.875 per 1 million output tokens. The appropriate choice depends on whether the application values lower cost or faster, more immediate processing. Context caching is supported, which can help applications that repeatedly refer to the same large context, although the supplied research does not specify the separate cache pricing details.

Search and Maps grounding may involve separate request-based charges after applicable monthly free allowances. Developers should therefore estimate both token costs and tool-related charges when calculating the total cost of a grounded workflow.

Main strengths and trade-offs

Gemini 3.8 Flash’s clearest strength is the combination of a very large context window, configurable reasoning, multimodal inputs, and agent-oriented tools. A single workflow can potentially combine a large repository or document set with visual or audio evidence, then use search, file search, function calling, or code execution to continue the task.

Its Flash positioning also makes it a candidate for high-volume applications that need more reasoning and tool support than a lightweight model provides. Batch and Flex options can improve cost efficiency for workloads that do not require the most immediate response path.

The main trade-off is that the model is not a universal media generator. It accepts many input formats but returns text only. It may also become more expensive or slower when high thinking effort produces more output tokens. Occasional latency and timeout issues are documented, and tool use, grounding, and computer use introduce additional implementation and reliability considerations.

The model’s documented knowledge cutoff is March 2026. Google cautions that knowledge in some domains may remain limited to January 2025, consistent with the Gemini 3 model family. Search grounding can provide newer external information during a request, but it does not change the model’s underlying knowledge cutoff.

When to choose Gemini 3.8 Flash

Choose Gemini 3.8 Flash when the workload benefits from a large working context and multi-step execution but still needs a Flash-oriented cost and latency profile. It is a strong candidate for:

  • Long-running software-engineering and repository-analysis tasks
  • Autonomous agents that call functions or use external tools
  • Large-document, PDF, and multimodal research workflows
  • Enterprise knowledge work involving search, file retrieval, or code execution
  • High-volume applications that need configurable reasoning
  • Applications that can use batch or Flex processing to reduce introductory token costs

Another type of model may be more appropriate when the priority is native image, audio, video, or speech generation; when the task is simple enough that a smaller and cheaper model is sufficient; or when the application cannot tolerate occasional timeouts and variable token usage. A larger reasoning model may also be preferable for tasks where maximum reasoning quality matters more than Flash-style speed and cost, although the supplied research does not identify a specific alternative or provide a direct price comparison.

Limitations to plan for

Gemini 3.8 Flash should not be treated as a guarantee of factual accuracy or autonomous correctness. Its responses can contain hallucinations, and the model may make errors even when it has a large context or access to tools. Applications should validate generated code, check retrieved information, and require human review for high-impact decisions.

Computer use is currently identified as a preview capability, so its behavior and availability may change. Feature support can also vary according to the API surface, account, region, enabled tools, and product configuration. Structured outputs are documented, but the supplied research does not verify a separate legacy JSON-mode capability.

Overall, Gemini 3.8 Flash is best understood as a text-generating, multimodal-input model for demanding agent and software workflows. Its one-million-token context, tool support, adjustable thinking, and comparatively low introductory token prices make it attractive for complex, high-volume applications, while its text-only output and variable reasoning cost limit its suitability for direct media generation or tightly predictable workloads.


Answers to Frequently Asked Questions

Who should use Gemini 3.8 Flash?
It is best suited to developers and organizations building long-running software-engineering agents, repository-analysis tools, enterprise research assistants, large-document workflows, and multimodal applications that use functions, search, file retrieval, code execution, or other external tools. Applications should still validate generated code and retrieved information because the model can hallucinate, experience occasional timeouts, and produce higher costs or latency at stronger thinking settings.
How much does Gemini 3.8 Flash cost?
For the paid Gemini API standard tier, introductory pricing through December 31, 2026 is $0.75 per 1 million input tokens and $3.75 per 1 million output tokens, including thinking tokens. From January 1, 2027, standard pricing is scheduled to increase to $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. Batch and Flex processing are available at half the introductory standard rates through December 31, 2026.
What input and output formats does Gemini 3.8 Flash support?
The model accepts text, images, video, audio, and PDF files as inputs. Its native output is text only, so it can analyze supplied media but does not directly generate images, audio, video, music, or speech.
What is Gemini 3.8 Flash?
Gemini 3.8 Flash is Google DeepMind’s generally available model for long-running software engineering, autonomous agents, multimodal analysis, and complex knowledge workflows. Its canonical model identifier is "gemini-3.8-flash", and it is available through the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, Gemini applications, and other Google channels.
How large is Gemini 3.8 Flash’s context window?
Gemini 3.8 Flash supports up to 1,048,576 input tokens, or roughly one million tokens, and has a maximum output capacity of 65,536 tokens. This makes it suitable for analyzing large code repositories, long documents, multiple research files, and extended agent workflows.


Sources 6
Provider

About Google DeepMind