What is Gemini 3.1 Flash Live Preview?
Gemini 3.1 Flash Live Preview is a multimodal Google DeepMind model designed for low-latency dialogue. Its canonical model identifier is gemini-3.1-flash-live-preview. Rather than treating every request as a separate text completion, developers use it through Google's bidirectional Gemini Live API to maintain an interactive session in which audio, text, images, and video can be exchanged in real time.
This makes the model particularly relevant to voice assistants, conversational customer-service systems, interactive audio products, and applications that combine speech with camera or video input. A user might speak to an application while sharing an image, or a voice agent might receive live audio and return spoken answers while also exposing a text transcript.
Google lists the model as a legacy preview release. It remains accessible according to the supplied documentation, but Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments. That positioning matters: Gemini 3.1 Flash Live Preview can still be useful for existing integrations and supported use cases, but it is not the provider's preferred starting point for new projects.
Modalities and live capabilities
The model accepts four input types: text, images, audio, and video. Its direct outputs are text and generated audio. Audio output enables spoken conversations, while text output can be used for transcripts, captions, application logic, or interfaces that display the agent's response.
- Input: text, images, audio, and video.
- Output: text and audio.
- Session type: real-time, bidirectional Live API sessions.
- Audio: generated speech for voice interactions.
- Streaming: supported through Live API sessions.
- Search: Google Search grounding is supported.
- Tools: synchronous function calling is supported.
In practical terms, a Live API application can send media to the model, receive streamed responses, and connect the model to application-controlled functions. Synchronous function calling means the model can request an external action, but it waits for the application to return the tool result before continuing its response.
The model does not support image generation, so it should not be selected when the same model must create or edit images. It also lacks structured outputs, which means it is not a good fit for workflows that require every response to conform to a provider-enforced JSON schema.
Technical specifications and limits
| Specification | Verified value |
|---|---|
| Model identifier | gemini-3.1-flash-live-preview |
| Provider | Google DeepMind |
| Release date | March 11, 2026 |
| Current status | Legacy preview; currently accessible in the supplied research |
| Input context limit | 131,072 tokens |
| Maximum output | 65,536 tokens |
| Live API | Supported |
| Streaming | Supported through Live API sessions |
The 131,072-token input context limit provides room for substantial conversation history and multimodal session context, although the usable amount in a live application will also depend on how much audio, video, text, and tool information the application sends. The maximum output limit is 65,536 tokens, but real-time voice responses will normally be much shorter because conversational latency is more important than producing very long answers.
Thinking controls and tool use
Gemini 3.1 Flash Live uses thinkingLevel rather than the older thinkingBudget configuration. The documented levels are minimal, low, medium, and high. The default is minimal, which is intended to reduce latency during interactive conversations.
These settings provide a way to trade response speed against additional reasoning effort. Minimal thinking is the natural starting point for a voice assistant that needs to respond quickly. Higher levels may be more appropriate when the conversation involves a more difficult decision or a tool-selection problem, but they can work against the immediate responsiveness expected from a live dialogue system.
Function calling is available, but only synchronously. For example, a voice agent can call an application function to look up an order or retrieve account information, then continue after the application sends the result. It cannot use asynchronous function calling to keep speaking while a long-running tool operation proceeds in the background. Google Search grounding is supported, while Google Maps grounding is not.
Pricing and cost trade-offs
Google's current Gemini Developer API pricing groups Gemini 3.1 Flash Live Preview with Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The supplied pricing is:
| Usage | Price |
|---|---|
| Text input | $0.75 per 1 million tokens |
| Audio input | $3.00 per 1 million tokens or $0.005 per audio minute |
| Image or video input | $1.00 per 1 million tokens or $0.002 per minute |
| Text output | $4.50 per 1 million tokens |
| Audio output | $12.00 per 1 million tokens or $0.018 per audio minute |
Google AI Studio provides a free tier subject to its applicable limits. The token and per-minute prices describe API usage; actual spending depends on session duration, the amount of audio or video transmitted, the number of responses, and the amount of generated speech.
The model's value proposition is therefore not simply a low price in every category. It is the combination of live interaction and relatively economical input processing. Audio output is priced higher than text output, so an application that can display text instead of speaking every response may reduce costs. Conversely, spoken output is essential for hands-free assistants and other voice-first products.
Main strengths and limitations
The most important strength is specialization for real-time interaction. A model that can receive several media types and return generated audio is more suitable for a live voice session than a text-only model connected to a separate speech pipeline. Live API streaming also supports a more immediate conversational experience.
- Low-latency design for interactive dialogue.
- Multimodal input covering text, images, audio, and video.
- Native audio generation for spoken responses.
- Synchronous function calling for connected applications.
- Google Search grounding for supported current-information workflows.
- Configurable thinking levels, including a latency-focused minimal setting.
- Large 131,072-token input context and 65,536-token maximum output.
Its limitations are equally important. It is a legacy preview model, and Google recommends Gemini 3.8 Live for most new deployments. It does not support asynchronous function calling, proactive audio, affective dialogue, structured outputs, code execution, file search, URL context, Google Maps grounding, batch API processing, caching, or image generation. Developers should verify current availability and migration guidance before building a long-lived integration around it.
Reasoning and coding suitability
The configurable thinking levels indicate that the model can apply different amounts of internal reasoning effort, but the supplied research does not provide a benchmark result or a provider-published reasoning score for this exact model. Its design priority is responsive live conversation, so it should not automatically be treated as the best choice for the most demanding analytical tasks.
The research database gives the model an editorial coding score of 5 out of 10 and a reasoning score of 7 out of 10. These are comparative editorial assessments, not Google-published benchmark results. The model can support application tools through function calling, but it does not support code execution. That makes it suitable for a voice interface that triggers predefined application actions, not for an environment where the model must run and test code as part of its answer.
Best use cases
Gemini 3.1 Flash Live Preview is a reasonable fit when immediate spoken interaction is central to the product:
- Real-time voice assistants and hands-free interfaces.
- Conversational customer-service agents that need controlled application functions.
- Interactive audio applications and spoken information services.
- Live sessions that combine microphone input with camera images or video.
- Voice interfaces that need Google Search grounding.
- Multimodal prototypes that need text transcripts alongside spoken responses.
For example, a support agent could hear a customer's question, inspect an uploaded image, call a synchronous order-status function, and speak the result. A visual assistant could receive camera input and answer questions aloud. These workflows use the model's live multimodal design directly rather than treating it as a general text model.
When to choose this model
Choose Gemini 3.1 Flash Live Preview when an existing or planned application specifically needs a low-latency Live API session, audio output, multimodal input, and straightforward synchronous tools. It can also be attractive when the application benefits from minimal thinking by default and the available pricing fits its audio and token usage.
Choose another option when the project is a new voice-agent deployment and long-term alignment with Google's recommended lineup is more important than compatibility with this preview model. Google specifically recommends Gemini 3.8 Live for most new low-latency voice-agent work. A different model or architecture may also be more appropriate if the application requires asynchronous tools, guaranteed structured JSON, code execution, image generation, Maps grounding, or batch processing.
The central trade-off is capability versus responsiveness and scope. Gemini 3.1 Flash Live Preview offers a focused set of features for real-time spoken interaction, rather than trying to cover every API workflow. It is strongest when fast multimodal conversation is the primary requirement and weaker when the application needs extensive orchestration, strict output schemas, or broader developer tooling.
Migration considerations
Developers migrating from Gemini 2.5 Flash Live should update the model identifier, replace thinking-budget configuration with thinking levels, and ensure that client code can process multiple content parts in a single server event. These changes are important for applications that assume a particular event structure or use the older configuration model.
Because the release is marked as legacy preview, teams should also treat availability and future support as planning considerations. The supplied research records no shutdown date, but the absence of a shutdown date is not a promise of indefinite support. For new systems, compare the required features and migration cost with Gemini 3.8 Live before committing to this model.

