What is Gemini 3.8 Live Extended Thinking?
Gemini 3.8 Live Extended Thinking is a stable Google model for real-time, bidirectional voice applications. It belongs to the Gemini 3.8 Audio family and is accessed through the Gemini Live API using the model identifier gemini-3.8-live-extended-thinking.
The model's defining feature is extended background reasoning. A conventional voice assistant often tries to produce a response as quickly as possible after the speaker finishes. This model is intended for requests that require several steps: understanding the request, planning a solution, consulting external information, calling one or more tools, and then delivering a useful spoken result.
It can remain active while that work continues. Applications may receive conversational audio, including progress-oriented filler responses, while the model reasons or waits for an external tool. This can make a long-running operation feel more natural than leaving the user in silence, although it also means that developers must manage the session lifecycle more carefully.
Where it fits in Google's lineup
Gemini 3.8 Live Extended Thinking is a specialized Live API model rather than a general-purpose text model or an image-generation model. Its focus is real-time, multimodal interaction with stronger reasoning for complex voice workflows. The supplied model documentation positions it separately from the standard Gemini 3.8 Live model, which is the more appropriate choice when the priority is the shortest possible response path for simple interactions.
This positioning matters because the model is not intended to maximize every capability at once. It supports live streaming, audio output, asynchronous tools, and Google Search grounding, but the supplied documentation does not list image generation, video generation, structured outputs, code execution, file search, Google Maps grounding, URL context, caching, or batch processing for this exact model.
Core capabilities and supported modalities
The model accepts four input types:
- Text
- Images
- Audio
- Video
It produces text and audio outputs. Audio output is particularly important for live voice agents, while text output can be useful for transcripts, application state, accessibility features, or interfaces that combine spoken and written responses. Although the model can process images and video as inputs, it should not be treated as an image or video creation system: image and video output are not listed as supported.
Live API streaming provides the bidirectional connection needed for an interactive session. The client can send user input and receive model events as the exchange develops, rather than waiting for a single conventional request-and-response transaction.
How extended thinking changes session management
The most important implementation detail is that a completed spoken turn does not necessarily mean the whole interaction is finished. A turnComplete event can indicate that the user's utterance has ended while background reasoning or tool execution continues.
Clients should therefore monitor interaction_status and wait until the status becomes IDLE before treating the workflow as complete. This distinction is easy to miss if an application assumes that every turn-completion event marks the end of all model activity.
For example, a user might ask a support agent to inspect a set of error codes and recommend a configuration change. The model may finish processing the spoken request, call an external diagnostic service, reason over the returned data, and then provide a final spoken explanation. The application needs to remain ready for those later events instead of closing the session immediately after the first completion signal.
Asynchronous function calling
Function calling allows the model to request that an application run an external operation, such as a search, database query, booking lookup, or diagnostic check. For Gemini 3.8 Live Extended Thinking, function calling is documented as asynchronous and non-blocking only. Synchronous blocking execution is not supported.
In practice, the client must acknowledge and manage tool requests without assuming that the model will pause in a simple, fully blocking state. The model may continue the conversational experience while the tool runs, including generating audio fillers. Developers should design for additional model messages after an initial turnComplete event and should keep tool results associated with the correct live interaction.
Reasoning capabilities
Thinking can be configured at low, medium, or high levels. Minimal thinking is not supported according to the supplied model information. The available levels provide a way to balance response effort against latency and cost, although the exact behavior of each level is application-dependent.
Extended Thinking is most useful when a request has multiple dependencies or requires planning before speaking with confidence. Examples include coordinating several travel constraints, troubleshooting a technical system, checking a multi-step programming solution, or using a business API that may take time to respond.
The benefit is not simply a longer answer. The important distinction is that reasoning can continue in the background while the live interaction remains active. The trade-off is that simple commands may not benefit from the additional reasoning path and may feel unnecessarily complex compared with a faster live model.
Technical specifications
| Specification | Verified value |
|---|---|
| Model ID | gemini-3.8-live-extended-thinking |
| Provider | Google DeepMind |
| Status | Stable and generally available |
| Context or input limit | 131,072 tokens |
| Maximum output | 65,536 tokens |
| Input modalities | Text, image, audio, and video |
| Output modalities | Text and audio |
| Live API | Supported |
| Thinking levels | Low, medium, and high |
| Function calling | Asynchronous, non-blocking only |
| Google Search grounding | Supported |
| Structured outputs | Not supported |
| Batch API | Not supported |
| Prompt caching | Not supported on the supplied model documentation |
The token limits are the documented maximums supplied for this model. They should not be interpreted as a guarantee that every live interaction will use the full context or output capacity, particularly when audio, streaming events, tool calls, and application-side state are also part of the workflow.
Pricing for text, audio, image, and video inputs
Google lists this model in the Live API pricing group. Paid standard pricing is charged according to the type of content processed:
| Usage type | Price |
|---|---|
| Text input | $0.75 per 1 million tokens |
| Audio input | $3.00 per 1 million tokens or $0.005 per minute |
| Image or video input | $1.00 per 1 million tokens or $0.002 per minute |
| Text output | $4.50 per 1 million tokens |
| Audio output | $12.00 per 1 million tokens or $0.018 per minute |
Audio output is therefore substantially more expensive per token than text output, while minute-based pricing may be easier to estimate for voice applications. Actual application cost depends on the amount and duration of input and output, the selected thinking level, the number of tool interactions, and the application's session design. The supplied research does not specify a separate subscription price or a guaranteed free allowance for this model.
Coding, tools, and practical integration
The model can be useful in coding-related voice workflows, such as a programming tutor that walks through a multi-step solution or a support agent that discusses logs and configuration files. The editorial coding assessment supplied for the model is 7 out of 10, but this is a comparative editorial score, not a provider-published benchmark.
Code execution itself is not listed as supported. That means the model may reason about code or request an external application tool, but developers should not assume that Google will execute arbitrary code inside this model's native capability set. Similarly, file search and URL context are not listed as supported, so any document retrieval or web-page processing workflow must use an explicitly supported integration rather than an assumed built-in feature.
Google Search grounding is listed as supported. This can help an application obtain current external information during a live interaction, but grounding does not change the model's underlying knowledge cutoff. No specific knowledge cutoff date is published for this exact model in the supplied research.
Main strengths and limitations
Strengths
- Designed for complex real-time voice interaction rather than only short command-and-response exchanges.
- Can continue reasoning after a spoken turn and can provide conversational audio while work is in progress.
- Supports multimodal inputs, including text, images, audio, and video.
- Supports asynchronous tool use for long-running or non-blocking workflows.
- Offers configurable low, medium, and high thinking levels.
- Supports Google Search grounding for current external information.
- Provides both text and audio outputs for applications that need voice plus transcript or interface content.
Limitations
- It is more demanding to integrate than a simple request-and-response or low-latency voice model.
turnCompletedoes not necessarily indicate that reasoning and tools have finished; clients must trackinteraction_status.- Function calling is asynchronous only; synchronous blocking tools are not supported.
- Structured outputs are not supported, and JSON mode should not be assumed.
- Image generation, video generation, code execution, file search, Maps grounding, URL context, caching, and batch processing are not listed for this exact model.
- Audio output carries a higher listed per-token price than text output.
- Simple, latency-sensitive commands may not justify the additional reasoning and session-management overhead.
Best use cases
Gemini 3.8 Live Extended Thinking is a strong fit when the user expects an active voice conversation but the task cannot be completed reliably in one immediate response. Suitable examples include:
- Technical support: inspect spoken descriptions, images, logs, or error information, then call diagnostic tools and explain the result.
- Travel and booking assistance: coordinate several searches or constraints while keeping the user informed during external requests.
- STEM and programming tutoring: work through multi-step reasoning and explain the process aloud.
- Operational assistants: interact with long-running business APIs without making the user wait in silence.
- Complex customer service: plan a response across account, policy, and service systems before delivering a final answer.
When to choose this model
Choose Gemini 3.8 Live Extended Thinking when background reasoning, asynchronous tools, and spoken progress are central requirements. It is particularly suitable when a live agent must remain conversational while external operations take place or when the task benefits from more deliberate planning.
Choose a simpler low-latency voice option, such as the standard Gemini 3.8 Live model referenced in the supplied research, when the application mostly handles short commands, quick questions, or predictable turn-taking. A text-focused model may be more appropriate for structured-output pipelines, while a model or workflow with code execution should be preferred when running code is a core requirement. The choice is therefore less about whether this model can produce a voice response and more about whether the application needs its additional reasoning lifecycle enough to justify higher complexity and potentially higher usage cost.
Bottom line
Gemini 3.8 Live Extended Thinking is a specialized real-time voice model for multi-step work. Its combination of multimodal input, audio output, configurable thinking, asynchronous function calling, and Search grounding makes it suitable for agents that need to investigate, plan, and act rather than simply answer immediately. The central implementation rule is to treat the interaction as complete only when the model reports an IDLE state. Teams that can manage that lifecycle will get the most value from the model; teams building fast, simple voice commands may be better served by a less reasoning-intensive alternative.

