Gemini 3.1

Gemini 3.1 Flash-Lite

by Google DeepMind · Deprecated; generally available and accessible until scheduled shutdown on May 7, 2027

Google’s Gemini 3.1 Flash-Lite is an efficiency-focused multimodal model for high-volume translation, classification, extraction, summarization, and lightweight agent workflows. It accepts text, images, video, audio, and PDFs, supports a 1-million-token context window and 65,536-token outputs, and offers tool use, structured outputs, caching, and discounted batch processing. It produces text only and is scheduled to shut down on May 7, 2027.

Text Reasoning Coding
Gemini 3.1 Flash-Lite is Google’s lightweight Gemini 3 model for applications where response speed, processing volume, and operating cost matter more than maximum reasoning depth. It combines multimodal input handling with text generation, tool use, structured outputs, and a very large context window. That makes it a practical choice for translation, classification, document extraction, summarization, routing, and repetitive agent steps. However, it is not designed for native image, audio, or video generation, and its documented shutdown date means teams should consider migration planning before adopting it for a long-lived system.
Outputs

What Gemini 3.1 Flash-Lite can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
10/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.1
Model type Lightweight
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff January 2025
Release date 2026-05-07
Status Deprecated; generally available and accessible until scheduled shutdown on May 7, 2027
Deprecation date 2026-05-07
Shutdown date 2027-05-07
Knowledge cutoff notes

Google's Gemini 3 developer guide lists January 2025 as the knowledge cutoff for Gemini 3.1 Flash-Lite. Search grounding and other external tools can provide newer information during use but do not change the underlying model cutoff.

Model notes

The canonical stable model ID is gemini-3.1-flash-lite. It accepts text, image, video, audio, and PDF inputs and returns text only. Google documents support for function calling, code execution, file search, Google Search grounding, Google Maps grounding, URL context, structured outputs, caching, batch inference, flex inference, and priority inference. The model supports configurable thinking levels. The earlier gemini-3.1-flash-lite-preview identifier was shut down on May 25, 2026. Google's deprecation page lists May 7, 2026 as the deprecation date for the stable model and May 7, 2027 as its shutdown date, with Gemini 3.5 Flash-Lite as the migration target. Editorial scores are comparative estimates, not vendor specifications.

Cost

Model pricing

Input $0.25 per 1M text/image/video tokens; $0.50 per 1M audio tokens; batch/flex: $0.125 per 1M text/image/video tokens and $0.25 per 1M audio tokens
Output $1.50 per 1M tokens standard; $0.75 per 1M tokens for batch/flex inference
Model guide

Gemini 3.1 Flash-Lite: Google’s Low-Cost Model for High-Volume Multimodal Work

Gemini 3.1 Flash-Lite is Google’s low-latency, cost-efficient multimodal model for high-volume workloads such as translation, classification, extraction, summarization, and lightweight tool-using agents. It accepts text, images, video, audio, and PDFs, supports a 1,048,576-token input context and 65,536-token outputs, and includes function calling, search grounding, structured outputs, caching, and batch processing. It remains accessible but is scheduled to shut down on May 7, 2027, with Gemini 3.5 Flash-Lite identified as the migration target.

What Gemini 3.1 Flash-Lite is

Gemini 3.1 Flash-Lite is a generally available model from Google, provided through the Gemini API and Google AI Studio, with related availability through Google’s enterprise AI platforms. Its canonical stable model identifier is gemini-3.1-flash-lite.

The model sits at the efficiency-focused end of Google’s Gemini 3 lineup. Rather than targeting the most demanding reasoning, research, or software-engineering tasks, it is intended to handle large numbers of relatively focused requests quickly and economically. Google describes this type of workload as including translation, simple data processing, and high-frequency agentic tasks.

In practical terms, Flash-Lite is best understood as a model for processing pipelines. A business might use it to classify incoming support tickets, extract fields from invoices, translate customer messages, summarize documents, or decide which downstream system should handle a request. It can also participate in tool-assisted workflows, but it should not automatically be treated as a substitute for a larger model on complex, long-horizon tasks.

Inputs, outputs, and context window

Gemini 3.1 Flash-Lite accepts text, images, video, audio, and PDF files. Its supported multimodal inputs allow an application to combine written instructions with visual documents, recordings, or other media. The model produces text output only: it does not natively generate images, audio, speech, or video.

SpecificationDocumented capability
Input typesText, images, video, audio, and PDFs
Maximum input context1,048,576 tokens
Maximum output65,536 tokens
Output typeText
Knowledge cutoffJanuary 2025

The one-million-token context limit is useful when a request depends on a long document or a large collection of related material. Context size is not the same as reasoning quality, however. A model may accept a large amount of information without being the best choice for interpreting complicated relationships across that information. For difficult research synthesis or complex planning, a more capable model may be more appropriate even if it costs more or responds more slowly.

The January 2025 knowledge cutoff applies to the underlying model. Search grounding and other external tools can supply newer information during a request, but they do not change the model’s internal training cutoff.

Tools and structured workflows

The model supports function calling, which allows it to request actions from software connected to the application. For example, it can select a customer-record lookup function, provide the required arguments, and let the application execute the operation. The model does not independently gain unrestricted access to a company’s systems; the application controls which functions exist and whether requested actions are actually performed.

Documented capabilities also include code execution, file search, URL context, Google Search grounding, Google Maps grounding, structured outputs, context caching, and several inference modes. Structured outputs are useful when the application needs predictable fields rather than free-form prose, such as a JSON-like result containing a ticket category, urgency level, language, and extracted reference number. The supplied research confirms structured-output support, but it does not establish a separate provider-defined “JSON mode” capability.

  • Function calling: Connects model responses to application-defined tools and operations.
  • Search grounding: Allows supported workflows to use Google Search or Google Maps information, with tool usage potentially incurring separate charges.
  • Code execution: Supports workflows that need programmatic computation.
  • File search and URL context: Helps incorporate information from supplied files or web URLs.
  • Structured outputs: Helps return information in an application-defined structure.
  • Context caching: Can reduce repeated processing costs for reused context, subject to storage charges.

These features make Gemini 3.1 Flash-Lite more than a simple text classifier, but its tool support does not remove the need for application-level safeguards. Developers still need to validate arguments, control permissions, handle failed calls, and check model-generated results before taking consequential actions.

Pricing and efficiency

Google’s documented standard pricing is $0.25 per 1 million input tokens for text, image, and video, $0.50 per 1 million audio input tokens, and $1.50 per 1 million output tokens, including thinking tokens. These rates make the model particularly relevant to systems that process many small or medium-sized requests.

Batch and flex processing reduce the listed rates. For batch or flex inference, text, image, and video input costs $0.125 per 1 million tokens, audio input costs $0.25 per 1 million tokens, and output costs $0.75 per 1 million tokens. These options are more suitable when immediate responses are less important than reducing cost.

Context caching is also supported. Standard cached-context pricing is documented as $0.025 per 1 million text, image, or video tokens and $0.05 per 1 million audio tokens, in addition to storage charges. Google Search grounding and Google Maps grounding may create separate tool-use charges, so the model’s token price should not be treated as the complete cost of every grounded request.

The prices above are provider-documented rates, not an estimate of total application cost. Actual spending also depends on input length, output length, repeated context, tool usage, traffic patterns, and the selected inference option.

Reasoning, coding, and speed trade-offs

Gemini 3.1 Flash-Lite supports configurable thinking levels, but it is positioned as an efficiency-oriented model rather than Google’s choice for the hardest reasoning problems. The supplied editorial assessment gives it a reasoning score of 7 out of 10, a coding score of 7 out of 10, a speed score of 10 out of 10, and a cost score of 10 out of 10. These are comparative editorial estimates, not scores published by Google and not benchmark results.

Its reasoning profile should be adequate for focused classification, extraction, transformation, summarization, and straightforward tool selection. Coding support can be useful for code generation, small transformations, data handling, and tool workflows. More demanding software engineering, complex debugging across a large codebase, autonomous planning, or research requiring several dependent decisions may benefit from a larger or more reasoning-focused model.

The central trade-off is straightforward: Flash-Lite exchanges some depth and sophistication for lower latency and lower cost. For a system processing thousands or millions of routine requests, that trade-off may be beneficial. For a small number of high-stakes requests where mistakes are expensive, the lower price may not justify additional review or orchestration.

Best use cases

Gemini 3.1 Flash-Lite is a strong candidate when requests are numerous, reasonably well-defined, and easy to validate. Suitable examples include:

  • Translating support tickets, reviews, messages, and other customer content.
  • Classifying tickets, documents, feedback, or moderation queues.
  • Extracting names, dates, totals, categories, and identifiers from PDFs, images, and other documents.
  • Summarizing routine reports, conversations, and submitted files.
  • Generating metadata, labels, routing decisions, or short descriptions.
  • Processing multimodal documents where the result can be represented as text or structured fields.
  • Handling repetitive customer-support steps with controlled function calls.
  • Performing lightweight agent tasks that use search, file retrieval, URL context, or application tools.

For these workloads, the large context window can be useful when each request includes substantial source material, while caching and batch processing can help reduce the cost of repeated or non-urgent work.

When to choose Gemini 3.1 Flash-Lite

Choose Gemini 3.1 Flash-Lite when throughput, latency, and token cost are primary requirements and the task can be broken into focused, testable steps. It is especially attractive when the application needs multimodal input but only text or structured text output.

A different model type may be preferable when the task requires deep reasoning, complex autonomous planning, advanced software engineering, or consistently high-quality research synthesis. A specialized generative model is also more appropriate when the application must create images, audio, speech, or video, because Flash-Lite returns text only.

Teams should also account for lifecycle status. Google’s deprecation documentation lists May 7, 2026 as the deprecation date and May 7, 2027 as the shutdown date for the stable identifier gemini-3.1-flash-lite. Google lists Gemini 3.5 Flash-Lite as the migration target. The earlier gemini-3.1-flash-lite-preview identifier was shut down on May 25, 2026 and should not be treated as the current stable model. Any new production integration should therefore isolate the model identifier and test the stated replacement before the shutdown deadline.

Limitations and final assessment

Gemini 3.1 Flash-Lite’s main limitation is not a lack of input flexibility; it is the boundary between efficient routine processing and demanding reasoning. It can accept many media types, use tools, and handle a very large context, but those features do not make it a native media-generation model or guarantee reliable performance on complex tasks.

It is most compelling as a fast, inexpensive processing layer for high-volume multimodal workloads. Its documented prices, large context limit, structured-output support, and tool integrations give developers several ways to build economical pipelines. Its scheduled shutdown, text-only output, January 2025 knowledge cutoff, and less ambitious reasoning profile are equally important when evaluating it. For short-lived or migration-ready systems, it can be a practical efficiency choice; for new long-term deployments, teams should compare it carefully with the listed successor and validate migration compatibility early.


Answers to Frequently Asked Questions

When will Gemini 3.1 Flash-Lite be discontinued?
Google lists May 7, 2026 as the deprecation date and May 7, 2027 as the shutdown date for the stable identifier gemini-3.1-flash-lite. Google lists Gemini 3.5 Flash-Lite as the migration target. The gemini-3.1-flash-lite-preview identifier was shut down on May 25, 2026.
Does Gemini 3.1 Flash-Lite support tools and structured outputs?
Yes. It supports function calling, code execution, file search, URL context, Google Search grounding, Google Maps grounding, structured outputs, context caching, and configurable thinking levels. Applications must still validate tool arguments, enforce permissions, and review model-generated results.
How much does Gemini 3.1 Flash-Lite cost?
Standard pricing is $0.25 per 1 million input tokens for text, image, and video, $0.50 per 1 million audio input tokens, and $1.50 per 1 million output tokens. Batch and flex processing reduce these rates to $0.125, $0.25, and $0.75 per 1 million tokens respectively. Tool usage, caching storage, and other services may add separate charges.
What is Gemini 3.1 Flash-Lite best used for?
Gemini 3.1 Flash-Lite is designed for high-volume, relatively focused tasks such as translation, classification, document and invoice data extraction, summarization, metadata generation, routing, and lightweight agent workflows.
What inputs and outputs does Gemini 3.1 Flash-Lite support?
It accepts text, images, video, audio, and PDF files, with a maximum input context of 1,048,576 tokens and a maximum output of 65,536 tokens. Its native output is text only; it does not generate images, audio, speech, or video.


Sources 6
Provider

About Google DeepMind