Gemini 3.8 Audio

Gemini 3.8 Live Extended Thinking

by Google DeepMind · Stable; generally available

A stable Google Live API model for complex, real-time voice interactions. It supports multimodal inputs, text and audio output, configurable background reasoning, asynchronous non-blocking function calls, conversational progress audio, and Google Search grounding, with documented 131,072-token input and 65,536-token output limits.

Text Speech Reasoning Coding
Gemini 3.8 Live Extended Thinking is designed for voice agents that need to do more than answer immediately. It can continue reasoning after a user has finished speaking, call external tools without blocking the conversation, and provide spoken progress or filler responses while a longer workflow is underway. That makes it a candidate for technical support, research, travel coordination, tutoring, and other applications where the agent must plan before completing a task. The model is available through Google's Gemini Live API and supports text, image, audio, and video inputs, with text and audio outputs.
Outputs

What Gemini 3.8 Live Extended Thinking can produce

Text Speech
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

9/10 Reasoning
7/10 Coding
7/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.8 Audio
Model type Reasoning
Context window 131K tokens
Maximum output 66K tokens
Release date September 15, 2026
Status Stable; generally available
Knowledge cutoff notes

Google's current model documentation does not publish a specific knowledge cutoff date for this exact Live Extended Thinking model. Live API Search grounding can provide current external information during use, but it does not establish or change the underlying model knowledge cutoff.

Model notes

The exact model ID is gemini-3.8-live-extended-thinking. It is a stable, generally available Live API model released on September 15, 2026. It supports configurable thinking levels of low, medium, and high; minimal thinking is not supported. Function calling is asynchronous only and requires non-blocking tool behavior. Clients should monitor interaction_status and wait for IDLE because turnComplete can occur before background reasoning and tool execution finish. Google lists Search grounding as supported, while code execution, file search, Google Maps grounding, image generation, structured outputs, URL context, caching, and batch API are not supported. Editorial scores are comparative estimates, not provider-published benchmarks.

Cost

Model pricing

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or $0.005 per audio minute; $1.00 per 1M image/video tokens or $0.002 per minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or $0.018 per audio minute
Model guide

Gemini 3.8 Live Extended Thinking for Voice Agents with Background Reasoning

Gemini 3.8 Live Extended Thinking is Google's stable Live API model for complex, real-time voice interactions. It combines audio-to-audio conversation with configurable background reasoning, asynchronous function calling, conversational progress updates, multimodal inputs, and Google Search grounding. Its main advantage is handling multi-step tasks without ending the interaction as soon as the user's spoken turn finishes; its main trade-off is greater implementation complexity and higher cost than a simpler low-latency voice model.

What is Gemini 3.8 Live Extended Thinking?

Gemini 3.8 Live Extended Thinking is a stable Google model for real-time, bidirectional voice applications. It belongs to the Gemini 3.8 Audio family and is accessed through the Gemini Live API using the model identifier gemini-3.8-live-extended-thinking.

The model's defining feature is extended background reasoning. A conventional voice assistant often tries to produce a response as quickly as possible after the speaker finishes. This model is intended for requests that require several steps: understanding the request, planning a solution, consulting external information, calling one or more tools, and then delivering a useful spoken result.

It can remain active while that work continues. Applications may receive conversational audio, including progress-oriented filler responses, while the model reasons or waits for an external tool. This can make a long-running operation feel more natural than leaving the user in silence, although it also means that developers must manage the session lifecycle more carefully.

Where it fits in Google's lineup

Gemini 3.8 Live Extended Thinking is a specialized Live API model rather than a general-purpose text model or an image-generation model. Its focus is real-time, multimodal interaction with stronger reasoning for complex voice workflows. The supplied model documentation positions it separately from the standard Gemini 3.8 Live model, which is the more appropriate choice when the priority is the shortest possible response path for simple interactions.

This positioning matters because the model is not intended to maximize every capability at once. It supports live streaming, audio output, asynchronous tools, and Google Search grounding, but the supplied documentation does not list image generation, video generation, structured outputs, code execution, file search, Google Maps grounding, URL context, caching, or batch processing for this exact model.

Core capabilities and supported modalities

The model accepts four input types:

  • Text
  • Images
  • Audio
  • Video

It produces text and audio outputs. Audio output is particularly important for live voice agents, while text output can be useful for transcripts, application state, accessibility features, or interfaces that combine spoken and written responses. Although the model can process images and video as inputs, it should not be treated as an image or video creation system: image and video output are not listed as supported.

Live API streaming provides the bidirectional connection needed for an interactive session. The client can send user input and receive model events as the exchange develops, rather than waiting for a single conventional request-and-response transaction.

How extended thinking changes session management

The most important implementation detail is that a completed spoken turn does not necessarily mean the whole interaction is finished. A turnComplete event can indicate that the user's utterance has ended while background reasoning or tool execution continues.

Clients should therefore monitor interaction_status and wait until the status becomes IDLE before treating the workflow as complete. This distinction is easy to miss if an application assumes that every turn-completion event marks the end of all model activity.

For example, a user might ask a support agent to inspect a set of error codes and recommend a configuration change. The model may finish processing the spoken request, call an external diagnostic service, reason over the returned data, and then provide a final spoken explanation. The application needs to remain ready for those later events instead of closing the session immediately after the first completion signal.

Asynchronous function calling

Function calling allows the model to request that an application run an external operation, such as a search, database query, booking lookup, or diagnostic check. For Gemini 3.8 Live Extended Thinking, function calling is documented as asynchronous and non-blocking only. Synchronous blocking execution is not supported.

In practice, the client must acknowledge and manage tool requests without assuming that the model will pause in a simple, fully blocking state. The model may continue the conversational experience while the tool runs, including generating audio fillers. Developers should design for additional model messages after an initial turnComplete event and should keep tool results associated with the correct live interaction.

Reasoning capabilities

Thinking can be configured at low, medium, or high levels. Minimal thinking is not supported according to the supplied model information. The available levels provide a way to balance response effort against latency and cost, although the exact behavior of each level is application-dependent.

Extended Thinking is most useful when a request has multiple dependencies or requires planning before speaking with confidence. Examples include coordinating several travel constraints, troubleshooting a technical system, checking a multi-step programming solution, or using a business API that may take time to respond.

The benefit is not simply a longer answer. The important distinction is that reasoning can continue in the background while the live interaction remains active. The trade-off is that simple commands may not benefit from the additional reasoning path and may feel unnecessarily complex compared with a faster live model.

Technical specifications

SpecificationVerified value
Model IDgemini-3.8-live-extended-thinking
ProviderGoogle DeepMind
StatusStable and generally available
Context or input limit131,072 tokens
Maximum output65,536 tokens
Input modalitiesText, image, audio, and video
Output modalitiesText and audio
Live APISupported
Thinking levelsLow, medium, and high
Function callingAsynchronous, non-blocking only
Google Search groundingSupported
Structured outputsNot supported
Batch APINot supported
Prompt cachingNot supported on the supplied model documentation

The token limits are the documented maximums supplied for this model. They should not be interpreted as a guarantee that every live interaction will use the full context or output capacity, particularly when audio, streaming events, tool calls, and application-side state are also part of the workflow.

Pricing for text, audio, image, and video inputs

Google lists this model in the Live API pricing group. Paid standard pricing is charged according to the type of content processed:

Usage typePrice
Text input$0.75 per 1 million tokens
Audio input$3.00 per 1 million tokens or $0.005 per minute
Image or video input$1.00 per 1 million tokens or $0.002 per minute
Text output$4.50 per 1 million tokens
Audio output$12.00 per 1 million tokens or $0.018 per minute

Audio output is therefore substantially more expensive per token than text output, while minute-based pricing may be easier to estimate for voice applications. Actual application cost depends on the amount and duration of input and output, the selected thinking level, the number of tool interactions, and the application's session design. The supplied research does not specify a separate subscription price or a guaranteed free allowance for this model.

Coding, tools, and practical integration

The model can be useful in coding-related voice workflows, such as a programming tutor that walks through a multi-step solution or a support agent that discusses logs and configuration files. The editorial coding assessment supplied for the model is 7 out of 10, but this is a comparative editorial score, not a provider-published benchmark.

Code execution itself is not listed as supported. That means the model may reason about code or request an external application tool, but developers should not assume that Google will execute arbitrary code inside this model's native capability set. Similarly, file search and URL context are not listed as supported, so any document retrieval or web-page processing workflow must use an explicitly supported integration rather than an assumed built-in feature.

Google Search grounding is listed as supported. This can help an application obtain current external information during a live interaction, but grounding does not change the model's underlying knowledge cutoff. No specific knowledge cutoff date is published for this exact model in the supplied research.

Main strengths and limitations

Strengths

  • Designed for complex real-time voice interaction rather than only short command-and-response exchanges.
  • Can continue reasoning after a spoken turn and can provide conversational audio while work is in progress.
  • Supports multimodal inputs, including text, images, audio, and video.
  • Supports asynchronous tool use for long-running or non-blocking workflows.
  • Offers configurable low, medium, and high thinking levels.
  • Supports Google Search grounding for current external information.
  • Provides both text and audio outputs for applications that need voice plus transcript or interface content.

Limitations

  • It is more demanding to integrate than a simple request-and-response or low-latency voice model.
  • turnComplete does not necessarily indicate that reasoning and tools have finished; clients must track interaction_status.
  • Function calling is asynchronous only; synchronous blocking tools are not supported.
  • Structured outputs are not supported, and JSON mode should not be assumed.
  • Image generation, video generation, code execution, file search, Maps grounding, URL context, caching, and batch processing are not listed for this exact model.
  • Audio output carries a higher listed per-token price than text output.
  • Simple, latency-sensitive commands may not justify the additional reasoning and session-management overhead.

Best use cases

Gemini 3.8 Live Extended Thinking is a strong fit when the user expects an active voice conversation but the task cannot be completed reliably in one immediate response. Suitable examples include:

  • Technical support: inspect spoken descriptions, images, logs, or error information, then call diagnostic tools and explain the result.
  • Travel and booking assistance: coordinate several searches or constraints while keeping the user informed during external requests.
  • STEM and programming tutoring: work through multi-step reasoning and explain the process aloud.
  • Operational assistants: interact with long-running business APIs without making the user wait in silence.
  • Complex customer service: plan a response across account, policy, and service systems before delivering a final answer.

When to choose this model

Choose Gemini 3.8 Live Extended Thinking when background reasoning, asynchronous tools, and spoken progress are central requirements. It is particularly suitable when a live agent must remain conversational while external operations take place or when the task benefits from more deliberate planning.

Choose a simpler low-latency voice option, such as the standard Gemini 3.8 Live model referenced in the supplied research, when the application mostly handles short commands, quick questions, or predictable turn-taking. A text-focused model may be more appropriate for structured-output pipelines, while a model or workflow with code execution should be preferred when running code is a core requirement. The choice is therefore less about whether this model can produce a voice response and more about whether the application needs its additional reasoning lifecycle enough to justify higher complexity and potentially higher usage cost.

Bottom line

Gemini 3.8 Live Extended Thinking is a specialized real-time voice model for multi-step work. Its combination of multimodal input, audio output, configurable thinking, asynchronous function calling, and Search grounding makes it suitable for agents that need to investigate, plan, and act rather than simply answer immediately. The central implementation rule is to treat the interaction as complete only when the model reports an IDLE state. Teams that can manage that lifecycle will get the most value from the model; teams building fast, simple voice commands may be better served by a less reasoning-intensive alternative.


Answers to Frequently Asked Questions

What input and output modalities does Gemini 3.8 Live Extended Thinking support?
The model accepts text, images, audio, and video as inputs and produces text and audio outputs. It supports live streaming, asynchronous function calling, and Google Search grounding, but image or video generation is not listed as supported.
When should you choose Gemini 3.8 Live Extended Thinking instead of a standard live voice model?
Choose Gemini 3.8 Live Extended Thinking for complex voice workflows that require background reasoning, multi-step planning, asynchronous tools, or long-running external operations. A standard Gemini 3.8 Live model is generally more suitable for simple commands and interactions where the shortest possible response time is the priority.
What is Gemini 3.8 Live Extended Thinking?
Gemini 3.8 Live Extended Thinking is a stable Google DeepMind model for real-time, bidirectional voice applications. It uses the Gemini Live API model identifier gemini-3.8-live-extended-thinking and is designed to continue reasoning, using tools, and responding conversationally during complex multi-step tasks.
How should developers know when a Gemini 3.8 Live Extended Thinking session is finished?
Developers should not treat a turnComplete event as the end of the entire interaction. Background reasoning or asynchronous tool execution may continue afterward. The client should monitor interaction_status and consider the workflow complete only when the status becomes IDLE.


Sources 8
Provider

About Google DeepMind