Gemini 3.5

Gemini 3.5 Flash-Lite

by Google DeepMind · General availability

Google's Gemini 3.5 Flash-Lite is a general-availability multimodal model optimized for high-throughput, low-latency workloads. It accepts text, images, video, audio, and PDFs, supports a 1-million-token context window and 65,536-token outputs, and provides tool use, search grounding, structured outputs, caching, batch processing, and computer use in preview.

Text Reasoning Coding
Gemini 3.5 Flash-Lite is the efficiency-focused member of Google's Gemini 3.5 family. It is designed for applications that need multimodal understanding, configurable reasoning, and tool use at high throughput while keeping inference costs and response latency relatively low.
Outputs

What Gemini 3.5 Flash-Lite can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
10/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.5
Model type Lightweight
Context window 1.05M tokens
Maximum output 66K tokens
Knowledge cutoff March 2026
Release date 2026-07-21
Status General availability
Knowledge cutoff notes

Google's model card lists March 2026 as the knowledge cutoff, while noting that some domains may have information limitations dating to January 2025 in line with the Gemini 3 model family. Search grounding can supply newer retrieved information without changing the underlying cutoff.

Model notes

The canonical Gemini API model ID is gemini-3.5-flash-lite. The model is based on Gemini 3.1 Flash-Lite and adds improved reasoning and agentic performance. Standard context caching is priced at $0.03 per 1 million cached input tokens, plus cache storage charges of $1.00 per 1 million tokens per hour. Batch and Flex input pricing is $0.15 per 1 million tokens and output pricing is $1.25 per 1 million tokens. Priority pricing is $0.54 input and $4.50 output per 1 million tokens. Computer use is supported in preview. Search grounding and Google Maps grounding are supported, with separate grounding charges after the applicable free allowance. The model produces text only and does not support image generation, audio generation, or the Live API.

Cost

Model pricing

Input $0.30 per 1 million tokens for standard text, image, video, and audio input; $0.15 per 1 million input tokens for Batch and Flex; $0.54 per 1 million input tokens for Priority.
Output $2.50 per 1 million output tokens for standard usage; $1.25 per 1 million output tokens for Batch and Flex; $4.50 per 1 million output tokens for Priority.
Model guide

Gemini 3.5 Flash-Lite: Google’s Low-Cost Model for High-Volume Multimodal Work

Gemini 3.5 Flash-Lite is Google's general-availability, low-latency multimodal model for high-volume agentic tasks, translation, document processing, classification, and data extraction. It accepts text, images, video, audio, and PDFs, supports a 1-million-token context window and 65,536-token outputs, and offers tool use, search grounding, structured outputs, caching, batch processing, and computer use in preview.

What is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is a general-availability multimodal model provided by Google. Its canonical Gemini API model identifier is gemini-3.5-flash-lite. The model is positioned for workloads where an application may need to process a large number of requests quickly and economically, rather than dedicate the highest available model capacity to every request.

Typical tasks include translation, classification, document parsing, structured data extraction, retrieval sub-agents, search-backed workflows, and other agentic subtasks. In this context, an agentic task is a step in which the model interprets information, decides what action or tool may be useful, and returns a result that can be used by a larger application workflow.

Google describes the model as being based on Gemini 3.1 Flash-Lite while adding stronger reasoning and agentic performance. That makes it a lightweight model in the Gemini lineup, but not merely a text classification system: it can process several media types, use supported tools, and produce structured responses for software systems.

Where it fits in the Gemini lineup

Gemini 3.5 Flash-Lite sits toward the speed- and cost-optimized end of Google's current model family. Its role is to handle high-volume or latency-sensitive work that may not require the maximum quality or depth available from a larger frontier model.

This positioning creates a practical trade-off. A larger model may be preferable for unusually difficult reasoning, complex coding, or high-stakes knowledge work. Gemini 3.5 Flash-Lite is more appropriate when the application needs to process many inputs, keep per-request costs predictable, or return results quickly. The supplied research does not provide a direct benchmark comparison against a larger sibling, so this distinction should be understood as product positioning and workload guidance rather than a measured ranking.

Inputs, outputs, and context limits

The model accepts text, images, video, audio, and PDF files. This allows a single workflow to combine written instructions with visual documents, recorded speech, video material, or scanned business records. For example, an application could submit a PDF invoice and ask the model to return selected fields in a structured format, or provide an image alongside text instructions for classification.

Gemini 3.5 Flash-Lite has a documented input context limit of 1,048,576 tokens, commonly described as a 1-million-token context window. A token is a unit of text or other model input used for processing; the exact number of tokens in a file depends on its content and encoding. The large context capacity is useful for long documents, collections of retrieved material, and multi-part multimodal tasks, although a large limit does not guarantee that every detail will be interpreted perfectly.

The maximum output is 65,536 tokens. The model produces text output only. It does not natively generate images, audio, or video, so it should not be selected for media-generation workflows. Its text output can nevertheless contain extracted fields, classifications, explanations, tool arguments, or other application-ready content.

Reasoning, coding, and tool support

Gemini 3.5 Flash-Lite supports thinking and configurable thinking levels. Thinking controls allow an application to balance reasoning effort against latency and cost, although the supplied research does not specify the exact behavior or settings for every use case. The model's documented positioning emphasizes improved reasoning and agentic performance compared with its Gemini 3.1 Flash-Lite basis.

Supported capabilities include function calling, code execution, search grounding, Google Maps grounding, file search, URL context, and structured outputs. Function calling allows the model to request an operation exposed by the application, such as looking up an account, checking an inventory system, or transforming data. The application—not the model—executes the function and returns the result. Search and Maps grounding can provide retrieved information during a workflow, with separate charges potentially applying after the relevant free allowance.

Structured outputs are useful when the response must follow a defined schema rather than arrive as free-form prose. This makes the model suitable for extracting fields from documents, assigning categories, or returning machine-readable results. The research confirms structured-output support but does not establish a separate, distinct JSON-mode capability; these should not automatically be treated as identical features.

Code execution is listed among the supported capabilities. The model also supports computer use in preview, which means the feature is not presented as a fully general-availability capability in the supplied documentation. The model does not support the Live API, making it unsuitable for applications that depend on that interface for continuous voice or live multimodal interaction.

Gemini 3.5 Flash-Lite pricing

Google's documented standard API pricing is:

Processing optionInput priceOutput price
Standard$0.30 per 1 million tokens$2.50 per 1 million tokens
Batch and Flex$0.15 per 1 million tokens$1.25 per 1 million tokens
Priority$0.54 per 1 million tokens$4.50 per 1 million tokens

Standard cached-input pricing is $0.03 per 1 million cached tokens, excluding cache storage charges. The supplied research lists cache storage at $1.00 per 1 million tokens per hour. Batch and Flex processing reduce the listed token rates and may be useful for jobs that do not require the lowest possible interactive latency. Priority processing has higher listed prices and is intended for workloads that place greater value on processing priority.

These are usage-based API prices rather than a single recurring subscription price. Actual costs depend on input volume, output volume, caching, processing mode, and any applicable grounding charges. Search grounding and Google Maps grounding may incur additional charges after the applicable free allowance.

Main strengths

  • High throughput: The model is specifically positioned for large request volumes and latency-sensitive applications.
  • Broad input coverage: Text, images, video, audio, and PDFs can be supplied as inputs.
  • Large context: The 1,048,576-token input limit supports long documents and substantial retrieved context.
  • Low listed token pricing: Standard pricing is substantially focused on cost-efficient processing, with lower Batch and Flex rates for suitable workloads.
  • Application integration: Function calling, code execution, search grounding, file search, URL context, and structured outputs support multi-step software workflows.
  • Flexible processing: Standard, Batch, Flex, Priority, and caching options allow applications to trade latency, cost, and processing requirements.
  • Text-based automation: Although the model does not generate media, its text output can represent extracted data, decisions, classifications, summaries, and tool requests.

Limitations and trade-offs

The most important limitation is that Gemini 3.5 Flash-Lite is text-output-only. It can understand images, audio, video, and PDFs, but it is not a native image, audio, or video generation model. Applications requiring generated media need a different model or service.

The model also does not support the Live API. That rules it out for use cases built around the documented Live API interface, particularly continuous interactive experiences. Computer use is available in preview, so teams should evaluate its behavior and operational suitability before relying on it for production-critical automation.

Google's model documentation notes that hallucinations, occasional latency or timeout issues, and uneven knowledge freshness remain possible. A large context window also should not be treated as a guarantee of perfect recall or accuracy across every supplied document. Applications that extract business data should validate important fields, handle failed or incomplete responses, and keep human review where mistakes would be costly.

The documented knowledge cutoff is March 2026. Google also notes that some domains may contain information limitations dating to January 2025, consistent with the broader Gemini 3 model family. Search grounding can retrieve newer information during use, but it does not change the model's underlying knowledge cutoff.

Best use cases

Gemini 3.5 Flash-Lite is a strong fit when an application needs to process many relatively independent tasks and return text or structured data quickly. Suitable examples include:

  • Extracting invoice, receipt, form, or contract fields from PDFs and images.
  • Classifying customer messages, documents, support tickets, or multimedia content.
  • Translating large volumes of text or multimodal material.
  • Summarizing long documents or preparing content for downstream systems.
  • Running retrieval sub-agents that search supplied files, URLs, or grounded sources.
  • Creating search-backed agents that need function calling and structured responses.
  • Processing audio or video inputs when the required result is a text transcript, classification, or analysis rather than generated media.
  • Providing rapid coding assistance or code-related analysis where the task does not require the deepest available reasoning.

Its cost structure can be especially useful for pipelines that can use Batch or Flex processing. Caching may also help when the same large context is reused across multiple requests, provided the storage and usage economics fit the workload.

When to choose Gemini 3.5 Flash-Lite

Choose Gemini 3.5 Flash-Lite when speed, volume, and cost efficiency are central requirements and the application benefits from multimodal input or tool use. It is particularly appropriate for document-heavy automation, classification, extraction, translation, and agentic subtasks that can be validated by application logic.

Consider a larger or more capable model instead when the main requirement is the highest quality on difficult reasoning, complex coding, or demanding knowledge work. Consider a media-generation model when the required output is an image, audio track, or video. A different interface is also more appropriate when the workflow depends on the Live API or production-ready computer-use behavior rather than preview functionality.

In short, Gemini 3.5 Flash-Lite is best understood as a fast processing component for multimodal and agentic systems. Its value comes from combining broad input support, a very large context window, tool integration, and low listed token rates—not from generating media or serving as the universal choice for the most difficult reasoning tasks.


Answers to Frequently Asked Questions

What are the main limitations of Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite produces text only, so it cannot natively generate images, audio, or video. It does not support the Live API, and computer use is still in preview. Applications should also account for possible hallucinations, latency or timeout issues, imperfect recall across large contexts, and the model's March 2026 knowledge cutoff.
Does Gemini 3.5 Flash-Lite support tools and structured outputs?
Yes. It supports function calling, code execution, search grounding, Google Maps grounding, file search, URL context, and structured outputs. Computer use is available in preview, while the Live API is not supported.
How much does Gemini 3.5 Flash-Lite cost?
Standard pricing is $0.30 per 1 million input tokens and $2.50 per 1 million output tokens. Batch and Flex pricing is $0.15 per 1 million input tokens and $1.25 per 1 million output tokens, while Priority pricing is $0.54 per 1 million input tokens and $4.50 per 1 million output tokens. Cached input is priced at $0.03 per 1 million tokens, excluding cache storage charges.
What is Gemini 3.5 Flash-Lite best used for?
Gemini 3.5 Flash-Lite is designed for high-volume, latency-sensitive workloads such as translation, classification, document parsing, structured data extraction, retrieval sub-agents, search-backed workflows, and other agentic tasks.
What types of input does Gemini 3.5 Flash-Lite support?
The model accepts text, images, video, audio, and PDF files. It supports an input context limit of 1,048,576 tokens, or approximately 1 million tokens, and can return up to 65,536 tokens of text output.


Sources 4
Provider

About Google DeepMind