Gemini 3.1

Gemini 3.1 Flash Live Preview

by Google DeepMind · Legacy preview; currently accessible; Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments

Gemini 3.1 Flash Live Preview is a legacy Google DeepMind model for low-latency multimodal voice applications. It supports text, image, audio, and video input, text and audio output, Live API streaming, Search grounding, configurable thinking levels, and synchronous function calling. The guide covers its technical limits, pricing, strengths, unsupported features, use cases, and reasons to consider Google's recommended Gemini 3.8 Live successor.

Text Speech Reasoning Coding
Gemini 3.1 Flash Live Preview is a Google DeepMind model built for live conversations and interactive voice agents rather than conventional one-shot text generation. It can process text, images, audio, and video during a Live API session, then respond with text or generated audio. The model is fast and relatively inexpensive for real-time interaction, but it is a legacy preview release with important exclusions, including structured outputs, asynchronous function calling, code execution, and image generation.
Outputs

What Gemini 3.1 Flash Live Preview can produce

Text Speech
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
5/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.1
Model type Multimodal
Context window 131K tokens
Maximum output 66K tokens
Release date March 11, 2026
Status Legacy preview; currently accessible; Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments
Knowledge cutoff notes

Google's model documentation does not provide a verified knowledge-cutoff date for this exact Live model.

Model notes

Canonical model ID is gemini-3.1-flash-live-preview. The model accepts text, images, audio, and video and returns text and audio. It supports Live API sessions, audio generation, synchronous function calling, Google Search grounding, and thinking levels minimal, low, medium, and high. Async function calling, proactive audio, affective dialogue, structured outputs, caching, batch API, code execution, file search, URL context, Google Maps grounding, and image generation are not supported. Pricing is shared with the Gemini 3.8 Live family on Google's current Gemini Developer API pricing page. No shutdown date has been announced.

Cost

Model pricing

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or $0.005 per audio minute; $1.00 per 1M image/video tokens or $0.002 per image/video minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or $0.018 per audio minute
Model guide

Gemini 3.1 Flash Live Preview: Low-Latency Voice AI for Real-Time Apps

Gemini 3.1 Flash Live Preview is Google's legacy preview model for real-time, multimodal voice applications. It accepts text, images, audio, and video, and produces text and audio through bidirectional Gemini Live API sessions. Its main advantages are low latency, audio generation, synchronous function calling, Search grounding, and configurable thinking levels. It remains accessible, but Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments.

What is Gemini 3.1 Flash Live Preview?

Gemini 3.1 Flash Live Preview is a multimodal Google DeepMind model designed for low-latency dialogue. Its canonical model identifier is gemini-3.1-flash-live-preview. Rather than treating every request as a separate text completion, developers use it through Google's bidirectional Gemini Live API to maintain an interactive session in which audio, text, images, and video can be exchanged in real time.

This makes the model particularly relevant to voice assistants, conversational customer-service systems, interactive audio products, and applications that combine speech with camera or video input. A user might speak to an application while sharing an image, or a voice agent might receive live audio and return spoken answers while also exposing a text transcript.

Google lists the model as a legacy preview release. It remains accessible according to the supplied documentation, but Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments. That positioning matters: Gemini 3.1 Flash Live Preview can still be useful for existing integrations and supported use cases, but it is not the provider's preferred starting point for new projects.

Modalities and live capabilities

The model accepts four input types: text, images, audio, and video. Its direct outputs are text and generated audio. Audio output enables spoken conversations, while text output can be used for transcripts, captions, application logic, or interfaces that display the agent's response.

  • Input: text, images, audio, and video.
  • Output: text and audio.
  • Session type: real-time, bidirectional Live API sessions.
  • Audio: generated speech for voice interactions.
  • Streaming: supported through Live API sessions.
  • Search: Google Search grounding is supported.
  • Tools: synchronous function calling is supported.

In practical terms, a Live API application can send media to the model, receive streamed responses, and connect the model to application-controlled functions. Synchronous function calling means the model can request an external action, but it waits for the application to return the tool result before continuing its response.

The model does not support image generation, so it should not be selected when the same model must create or edit images. It also lacks structured outputs, which means it is not a good fit for workflows that require every response to conform to a provider-enforced JSON schema.

Technical specifications and limits

SpecificationVerified value
Model identifiergemini-3.1-flash-live-preview
ProviderGoogle DeepMind
Release dateMarch 11, 2026
Current statusLegacy preview; currently accessible in the supplied research
Input context limit131,072 tokens
Maximum output65,536 tokens
Live APISupported
StreamingSupported through Live API sessions

The 131,072-token input context limit provides room for substantial conversation history and multimodal session context, although the usable amount in a live application will also depend on how much audio, video, text, and tool information the application sends. The maximum output limit is 65,536 tokens, but real-time voice responses will normally be much shorter because conversational latency is more important than producing very long answers.

Thinking controls and tool use

Gemini 3.1 Flash Live uses thinkingLevel rather than the older thinkingBudget configuration. The documented levels are minimal, low, medium, and high. The default is minimal, which is intended to reduce latency during interactive conversations.

These settings provide a way to trade response speed against additional reasoning effort. Minimal thinking is the natural starting point for a voice assistant that needs to respond quickly. Higher levels may be more appropriate when the conversation involves a more difficult decision or a tool-selection problem, but they can work against the immediate responsiveness expected from a live dialogue system.

Function calling is available, but only synchronously. For example, a voice agent can call an application function to look up an order or retrieve account information, then continue after the application sends the result. It cannot use asynchronous function calling to keep speaking while a long-running tool operation proceeds in the background. Google Search grounding is supported, while Google Maps grounding is not.

Pricing and cost trade-offs

Google's current Gemini Developer API pricing groups Gemini 3.1 Flash Live Preview with Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The supplied pricing is:

UsagePrice
Text input$0.75 per 1 million tokens
Audio input$3.00 per 1 million tokens or $0.005 per audio minute
Image or video input$1.00 per 1 million tokens or $0.002 per minute
Text output$4.50 per 1 million tokens
Audio output$12.00 per 1 million tokens or $0.018 per audio minute

Google AI Studio provides a free tier subject to its applicable limits. The token and per-minute prices describe API usage; actual spending depends on session duration, the amount of audio or video transmitted, the number of responses, and the amount of generated speech.

The model's value proposition is therefore not simply a low price in every category. It is the combination of live interaction and relatively economical input processing. Audio output is priced higher than text output, so an application that can display text instead of speaking every response may reduce costs. Conversely, spoken output is essential for hands-free assistants and other voice-first products.

Main strengths and limitations

The most important strength is specialization for real-time interaction. A model that can receive several media types and return generated audio is more suitable for a live voice session than a text-only model connected to a separate speech pipeline. Live API streaming also supports a more immediate conversational experience.

  • Low-latency design for interactive dialogue.
  • Multimodal input covering text, images, audio, and video.
  • Native audio generation for spoken responses.
  • Synchronous function calling for connected applications.
  • Google Search grounding for supported current-information workflows.
  • Configurable thinking levels, including a latency-focused minimal setting.
  • Large 131,072-token input context and 65,536-token maximum output.

Its limitations are equally important. It is a legacy preview model, and Google recommends Gemini 3.8 Live for most new deployments. It does not support asynchronous function calling, proactive audio, affective dialogue, structured outputs, code execution, file search, URL context, Google Maps grounding, batch API processing, caching, or image generation. Developers should verify current availability and migration guidance before building a long-lived integration around it.

Reasoning and coding suitability

The configurable thinking levels indicate that the model can apply different amounts of internal reasoning effort, but the supplied research does not provide a benchmark result or a provider-published reasoning score for this exact model. Its design priority is responsive live conversation, so it should not automatically be treated as the best choice for the most demanding analytical tasks.

The research database gives the model an editorial coding score of 5 out of 10 and a reasoning score of 7 out of 10. These are comparative editorial assessments, not Google-published benchmark results. The model can support application tools through function calling, but it does not support code execution. That makes it suitable for a voice interface that triggers predefined application actions, not for an environment where the model must run and test code as part of its answer.

Best use cases

Gemini 3.1 Flash Live Preview is a reasonable fit when immediate spoken interaction is central to the product:

  • Real-time voice assistants and hands-free interfaces.
  • Conversational customer-service agents that need controlled application functions.
  • Interactive audio applications and spoken information services.
  • Live sessions that combine microphone input with camera images or video.
  • Voice interfaces that need Google Search grounding.
  • Multimodal prototypes that need text transcripts alongside spoken responses.

For example, a support agent could hear a customer's question, inspect an uploaded image, call a synchronous order-status function, and speak the result. A visual assistant could receive camera input and answer questions aloud. These workflows use the model's live multimodal design directly rather than treating it as a general text model.

When to choose this model

Choose Gemini 3.1 Flash Live Preview when an existing or planned application specifically needs a low-latency Live API session, audio output, multimodal input, and straightforward synchronous tools. It can also be attractive when the application benefits from minimal thinking by default and the available pricing fits its audio and token usage.

Choose another option when the project is a new voice-agent deployment and long-term alignment with Google's recommended lineup is more important than compatibility with this preview model. Google specifically recommends Gemini 3.8 Live for most new low-latency voice-agent work. A different model or architecture may also be more appropriate if the application requires asynchronous tools, guaranteed structured JSON, code execution, image generation, Maps grounding, or batch processing.

The central trade-off is capability versus responsiveness and scope. Gemini 3.1 Flash Live Preview offers a focused set of features for real-time spoken interaction, rather than trying to cover every API workflow. It is strongest when fast multimodal conversation is the primary requirement and weaker when the application needs extensive orchestration, strict output schemas, or broader developer tooling.

Migration considerations

Developers migrating from Gemini 2.5 Flash Live should update the model identifier, replace thinking-budget configuration with thinking levels, and ensure that client code can process multiple content parts in a single server event. These changes are important for applications that assume a particular event structure or use the older configuration model.

Because the release is marked as legacy preview, teams should also treat availability and future support as planning considerations. The supplied research records no shutdown date, but the absence of a shutdown date is not a promise of indefinite support. For new systems, compare the required features and migration cost with Gemini 3.8 Live before committing to this model.


Answers to Frequently Asked Questions

Should new projects choose Gemini 3.1 Flash Live Preview?
Gemini 3.1 Flash Live Preview can be appropriate for existing integrations or applications that need low-latency multimodal voice interaction, audio output, and synchronous tools. However, it is classified as a legacy preview release, and Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments.
How much does Gemini 3.1 Flash Live Preview cost?
The listed prices are $0.75 per 1 million text input tokens, $3.00 per 1 million audio input tokens or $0.005 per audio minute, $1.00 per 1 million image or video input tokens or $0.002 per minute, $4.50 per 1 million text output tokens, and $12.00 per 1 million audio output tokens or $0.018 per audio minute. Google AI Studio also provides a free tier subject to applicable limits.
What inputs and outputs does Gemini 3.1 Flash Live Preview support?
The model accepts text, images, audio, and video as inputs. It generates text and audio outputs, allowing applications to provide spoken responses alongside transcripts, captions, or text-based application logic.
Does Gemini 3.1 Flash Live Preview support function calling and Google Search grounding?
Yes. It supports synchronous function calling and Google Search grounding. With synchronous function calling, the model waits for the application to return a tool result before continuing its response. Asynchronous function calling and Google Maps grounding are not supported.
What is Gemini 3.1 Flash Live Preview used for?
Gemini 3.1 Flash Live Preview is designed for low-latency, real-time conversations through Google's bidirectional Gemini Live API. It is suitable for voice assistants, customer-service agents, interactive audio applications, and multimodal experiences that combine speech with images or video.


Sources 5
Provider

About Google DeepMind