Gemini 3.8

Gemini 3.8 Live

by Google DeepMind · Stable; generally available

Gemini 3.8 Live is Google’s stable Live API model for low-latency voice agents and real-time multimodal dialogue. It accepts text, images, audio, and video, produces text or native audio, and supports streaming, Search grounding, asynchronous function calling, interleaved reasoning, and proactive audio. It is less suitable for structured JSON, batch processing, prompt caching, image generation, or code execution.

Text Speech Reasoning Coding
Gemini 3.8 Live is Google’s stable Live API model for building voice agents and other applications that need rapid, natural interaction. It can listen to audio, inspect images or video, use text, call external tools, and answer with streamed text or native audio. The model’s main advantage is low-latency audio-to-audio conversation, while its main trade-offs are the lack of structured outputs, batch processing, prompt caching, image generation, and code execution.
Outputs

What Gemini 3.8 Live can produce

Text Speech
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
6/10 Coding
10/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.8
Model type Multimodal
Context window 131K tokens
Maximum output 66K tokens
Release date 2026-09-15
Status Stable; generally available
Knowledge cutoff notes

Google's current model documentation does not specify a knowledge cutoff for Gemini 3.8 Live. Search grounding can provide current external information during supported API sessions, but it does not change the underlying model cutoff.

Model notes

Gemini 3.8 Live is the default stable Live API model for most low-latency voice-agent experiences. It supports interleaved reasoning, asynchronous function calling, full-session client content updates, proactive audio, and Google Search grounding. Audio is the supported live response modality; applications requiring transcripts should enable output audio transcription. Proactive audio is permanently enabled. The model does not support thinking_level or thinking_config, structured outputs, prompt caching, batch processing, code execution, file search, URL context, Google Maps grounding, or image generation. Pricing includes separate token and approximate per-minute audio rates. Editorial scores are comparative estimates, not vendor-provided benchmarks.

Cost

Model pricing

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or approximately $0.005 per minute; $1.00 per 1M image/video tokens or approximately $0.002 per minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or approximately $0.018 per minute
Model guide

Gemini 3.8 Live: Google’s Stable Real-Time Voice Model

Gemini 3.8 Live is Google’s stable, low-latency model for real-time voice agents and live multimodal dialogue. It accepts text, images, audio, and video, and produces text or native audio. Its Live API support includes bidirectional streaming, asynchronous function calling, Search grounding, interleaved reasoning, proactive audio, and session content updates. It is optimized for responsive conversations rather than batch processing, strict structured output, image generation, or workloads that depend on prompt caching.

What is Gemini 3.8 Live?

Gemini 3.8 Live is a Google model designed for live, interactive conversations rather than conventional one-request-at-a-time text generation. It is available through Google’s Gemini Live API and is described in the supplied Google documentation as the default stable model for most low-latency voice-agent experiences.

In practical terms, an application can maintain a bidirectional streaming session with the model. The user can speak, send text, provide an image, or share video, while the model processes the session and responds with text or generated audio. This makes it suitable for assistants that need to react quickly instead of waiting for a long reasoning process to finish before speaking.

Gemini 3.8 Live is provided by Google and belongs to the Gemini 3.8 model family. Its documented model identifier is gemini-3.8-live, and its status is stable and generally available.

Where it fits in Google’s model lineup

Gemini 3.8 Live occupies the real-time conversational position in Google’s current model catalog. It is not primarily a general-purpose batch model, an image-generation model, or a structured data extraction model. Its design prioritizes rapid live interaction, native audio, and the operational features needed by voice agents.

Google positions it as the recommended successor to the legacy gemini-3.1-flash-live-preview model. Migration involves changing the model identifier and reviewing Live API settings. In particular, Gemini 3.8 Live does not support thinking_level or thinking_config, so applications that used those settings with the older model should remove them.

The model is intended for applications where conversational responsiveness matters more than extended background reasoning. Its interleaved reasoning capability can support more considered responses, but the model remains specialized for live sessions rather than offline workloads.

Supported inputs and outputs

Gemini 3.8 Live accepts four input types:

  • Text
  • Images
  • Audio
  • Video

It produces text and audio outputs. Native audio output is central to its purpose: the model can participate in an audio-to-audio conversation without requiring an application to convert every user utterance into text and then convert the response back into speech through a separate pipeline.

For applications that need a written record of spoken responses, Google’s documentation indicates that output audio transcription should be enabled explicitly. Audio is the supported live response modality for generation, while text can also be returned as an output modality where the application requires it.

Gemini 3.8 Live does not generate images. Its multimodal capability therefore concerns understanding incoming media and producing text or speech, not creating visual assets.

Technical specifications

SpecificationGemini 3.8 Live
ProviderGoogle
Model identifiergemini-3.8-live
StatusStable; generally available
Input modalitiesText, image, audio, and video
Output modalitiesText and audio
Input token limit131,072 tokens
Maximum output65,536 tokens
Live API streamingSupported
Function callingSupported; asynchronous execution is the default
Search groundingSupported
Structured outputsNot supported
Prompt cachingNot supported
Batch APINot supported
Image generationNot supported

The 131,072-token input limit gives a session room for substantial conversational and multimodal context, while the documented 65,536-token output limit is considerably larger than what most spoken interactions will require. These limits should not be confused with guaranteed response length: live audio applications generally constrain responses through their session design, turn handling, and user experience requirements.

How its live conversation features work

The model’s main distinction is its session behavior. Live API sessions support bidirectional streaming, so the client can send content while the model is responding. This is more suitable for interruption-based conversation than a conventional request-and-response endpoint.

Asynchronous function calling is enabled by default. A function call is a request for the application to perform an external action, such as looking up an account record or retrieving information from another system. With asynchronous execution, the tool can run without unnecessarily blocking the conversation. Developers can use blocking execution when the model must wait for the result before continuing.

Search grounding is also supported. This allows an application to connect responses to current information retrieved through Google Search, although grounding does not change the model’s underlying knowledge cutoff. The supplied documentation does not specify a fixed knowledge cutoff for Gemini 3.8 Live.

Client content updates can be sent throughout a session with explicit user or model roles. An update with turn_complete=true interrupts active model generation. Content sent without completing the turn allows the server to wait for additional input. This distinction is important for applications that support interruptions, partial speech, corrections, or multiple pieces of context in one turn.

Proactive audio is permanently enabled for this model. Applications should therefore account for the possibility that the model can continue listening and produce audio according to the Live API session configuration. Proactive audio cannot be disabled, which may make the model less suitable for products that require tightly controlled, strictly turn-based speech behavior.

Gemini 3.8 Live pricing

Google lists separate rates for text, audio, and image or video tokens. The documented paid-tier input rates are $0.75 per million text tokens, $3.00 per million audio tokens, and $1.00 per million image or video tokens. Google also gives approximate audio and visual rates of $0.005 per minute for audio input and $0.002 per minute for image or video input.

Output pricing is $4.50 per million text tokens and $12.00 per million audio tokens. The approximate generated-audio rate is $0.018 per minute. These per-minute figures are useful for estimating voice workloads, but actual billing depends on the provider’s token accounting and the amount of content processed.

A free tier is available subject to applicable usage limits. Search grounding includes 5,000 free search requests per month shared across Gemini 3.x models. Additional Search grounding requests are priced at $14 per 1,000 queries according to the supplied pricing research.

The cost trade-off is straightforward: audio output costs more per token than text output, but native audio can remove the need for a separate speech-generation stage and can reduce interaction latency. Teams should estimate both model usage and external tool or search activity when calculating the total cost of a voice agent.

Main strengths

  • Low-latency voice interaction: The model is specifically designed for live dialogue and audio-to-audio applications.
  • Native audio: It can receive and generate audio directly rather than relying entirely on a text-only intermediary.
  • Broad multimodal input: An agent can combine spoken conversation with text, images, and video.
  • Responsive tool use: Asynchronous function calling lets external operations run without automatically freezing the conversation.
  • Current-information support: Search grounding is available for applications that need web-connected answers.
  • Stable availability: The model is generally available rather than being limited to preview status.
  • Large session capacity: The documented input and output limits support substantial conversational context and application instructions.

Main limitations

Gemini 3.8 Live’s specialization also defines its limits. It does not support structured outputs, so it is not a natural choice when every response must conform to a strict JSON schema. Applications can request text and use their own parsing or validation, but the model does not provide the documented structured-output capability.

Prompt caching and batch processing are not supported. This makes the model less appropriate for large offline workloads with repeated prefixes, scheduled processing, or cost-sensitive bulk inference. It also has no built-in code execution, file search, URL context, or Google Maps grounding in the supplied capability information.

Image generation is unavailable, and the model cannot be used as a single-model solution for a workflow that needs both live voice conversation and visual asset creation. Proactive audio is permanently enabled, which can complicate applications that require a fully manual push-to-talk interaction pattern.

Although the model supports interleaved reasoning, it is optimized for responsiveness rather than for long, deliberative research tasks. The supplied editorial assessment rates its reasoning at 7 out of 10 and coding at 6 out of 10, but these are comparative editorial scores, not Google-published benchmarks. Coding assistance is possible through text interaction and tool calls, but the model has no built-in code execution capability.

When to choose Gemini 3.8 Live

Choose Gemini 3.8 Live when the central requirement is a fast, natural conversation with streamed audio. Suitable examples include:

  • Customer-service and support voice agents
  • Voice-controlled applications
  • Interactive tutoring and language practice
  • Multimodal assistants that can discuss images or video during a conversation
  • Field-support agents that combine spoken instructions with visual context
  • Tool-using assistants that need to call business systems while maintaining dialogue

It is particularly attractive when a separate speech-recognition and text-to-speech pipeline would introduce unwanted delay or operational complexity. Search grounding can also help when the assistant must answer questions involving current external information.

When another option may be more appropriate

Use a different model or workflow when strict machine-readable output is the priority. Gemini 3.8 Live does not support structured outputs, so a model with native schema-constrained generation is a better fit for dependable JSON extraction or automated data pipelines.

A batch-oriented model is more appropriate for offline document processing, scheduled jobs, or high-volume inference. A model with prompt caching may be more economical when the same large instructions or context are sent repeatedly. Applications that need image generation should use a model designed to produce images rather than treating Gemini 3.8 Live as a general media-generation system.

For highly deliberative research or complex coding workflows, a model optimized for extended reasoning, code execution, or repository tools may offer a better capability trade-off even if it responds more slowly. Gemini 3.8 Live is the stronger choice when immediacy and spoken interaction outweigh those requirements.

Migration and design checklist

Teams moving from the older gemini-3.1-flash-live-preview model should change the model identifier to gemini-3.8-live and review session configuration rather than assuming complete behavioral compatibility.

  • Remove unsupported thinking_level and thinking_config settings.
  • Review asynchronous function-calling behavior and decide whether any tool must block the conversation.
  • Design for permanently enabled proactive audio.
  • Handle interruptions and partial updates using the Live API turn-completion rules.
  • Enable output audio transcription when written transcripts are required.
  • Do not depend on structured outputs, prompt caching, batch processing, or code execution.
  • Estimate audio token costs separately from text, image, and video usage.

Overall, Gemini 3.8 Live is best understood as a stable real-time conversation engine. Its value comes from combining low-latency native audio, multimodal input, streaming sessions, search grounding, and responsive tool use. Its design is less suitable for strict schemas, offline throughput, visual generation, or workflows where extended reasoning and execution tools matter more than conversational speed.


Answers to Frequently Asked Questions

What are the limitations of Gemini 3.8 Live?
Gemini 3.8 Live does not support structured outputs, prompt caching, batch processing, image generation, code execution, file search, URL context, or Google Maps grounding. It also has permanently enabled proactive audio and does not support thinking_level or thinking_config settings.
How much does Gemini 3.8 Live cost?
Google’s documented paid-tier rates are $0.75 per million text input tokens, $3.00 per million audio input tokens, and $1.00 per million image or video input tokens. Output costs are $4.50 per million text tokens and $12.00 per million audio tokens. Approximate audio rates are $0.005 per minute for input and $0.018 per minute for generated audio.
What are the main capabilities of Gemini 3.8 Live?
Gemini 3.8 Live supports low-latency bidirectional streaming, native audio conversations, multimodal input, asynchronous function calling, search grounding, and interleaved reasoning. It is designed primarily for voice agents and other live conversational applications.
What is Gemini 3.8 Live?
Gemini 3.8 Live is Google’s stable real-time voice model for interactive, bidirectional conversations. Through the Gemini Live API, it can accept text, images, audio, and video, and generate text or audio responses.
What is the model identifier for Gemini 3.8 Live?
The documented model identifier is gemini-3.8-live. It is generally available and positioned as the successor to the legacy gemini-3.1-flash-live-preview model.


Sources 8
Provider

About Google DeepMind