What is Gemini 3.8 Live?
Gemini 3.8 Live is a Google model designed for live, interactive conversations rather than conventional one-request-at-a-time text generation. It is available through Google’s Gemini Live API and is described in the supplied Google documentation as the default stable model for most low-latency voice-agent experiences.
In practical terms, an application can maintain a bidirectional streaming session with the model. The user can speak, send text, provide an image, or share video, while the model processes the session and responds with text or generated audio. This makes it suitable for assistants that need to react quickly instead of waiting for a long reasoning process to finish before speaking.
Gemini 3.8 Live is provided by Google and belongs to the Gemini 3.8 model family. Its documented model identifier is gemini-3.8-live, and its status is stable and generally available.
Where it fits in Google’s model lineup
Gemini 3.8 Live occupies the real-time conversational position in Google’s current model catalog. It is not primarily a general-purpose batch model, an image-generation model, or a structured data extraction model. Its design prioritizes rapid live interaction, native audio, and the operational features needed by voice agents.
Google positions it as the recommended successor to the legacy gemini-3.1-flash-live-preview model. Migration involves changing the model identifier and reviewing Live API settings. In particular, Gemini 3.8 Live does not support thinking_level or thinking_config, so applications that used those settings with the older model should remove them.
The model is intended for applications where conversational responsiveness matters more than extended background reasoning. Its interleaved reasoning capability can support more considered responses, but the model remains specialized for live sessions rather than offline workloads.
Supported inputs and outputs
Gemini 3.8 Live accepts four input types:
- Text
- Images
- Audio
- Video
It produces text and audio outputs. Native audio output is central to its purpose: the model can participate in an audio-to-audio conversation without requiring an application to convert every user utterance into text and then convert the response back into speech through a separate pipeline.
For applications that need a written record of spoken responses, Google’s documentation indicates that output audio transcription should be enabled explicitly. Audio is the supported live response modality for generation, while text can also be returned as an output modality where the application requires it.
Gemini 3.8 Live does not generate images. Its multimodal capability therefore concerns understanding incoming media and producing text or speech, not creating visual assets.
Technical specifications
| Specification | Gemini 3.8 Live |
|---|---|
| Provider | |
| Model identifier | gemini-3.8-live |
| Status | Stable; generally available |
| Input modalities | Text, image, audio, and video |
| Output modalities | Text and audio |
| Input token limit | 131,072 tokens |
| Maximum output | 65,536 tokens |
| Live API streaming | Supported |
| Function calling | Supported; asynchronous execution is the default |
| Search grounding | Supported |
| Structured outputs | Not supported |
| Prompt caching | Not supported |
| Batch API | Not supported |
| Image generation | Not supported |
The 131,072-token input limit gives a session room for substantial conversational and multimodal context, while the documented 65,536-token output limit is considerably larger than what most spoken interactions will require. These limits should not be confused with guaranteed response length: live audio applications generally constrain responses through their session design, turn handling, and user experience requirements.
How its live conversation features work
The model’s main distinction is its session behavior. Live API sessions support bidirectional streaming, so the client can send content while the model is responding. This is more suitable for interruption-based conversation than a conventional request-and-response endpoint.
Asynchronous function calling is enabled by default. A function call is a request for the application to perform an external action, such as looking up an account record or retrieving information from another system. With asynchronous execution, the tool can run without unnecessarily blocking the conversation. Developers can use blocking execution when the model must wait for the result before continuing.
Search grounding is also supported. This allows an application to connect responses to current information retrieved through Google Search, although grounding does not change the model’s underlying knowledge cutoff. The supplied documentation does not specify a fixed knowledge cutoff for Gemini 3.8 Live.
Client content updates can be sent throughout a session with explicit user or model roles. An update with turn_complete=true interrupts active model generation. Content sent without completing the turn allows the server to wait for additional input. This distinction is important for applications that support interruptions, partial speech, corrections, or multiple pieces of context in one turn.
Proactive audio is permanently enabled for this model. Applications should therefore account for the possibility that the model can continue listening and produce audio according to the Live API session configuration. Proactive audio cannot be disabled, which may make the model less suitable for products that require tightly controlled, strictly turn-based speech behavior.
Gemini 3.8 Live pricing
Google lists separate rates for text, audio, and image or video tokens. The documented paid-tier input rates are $0.75 per million text tokens, $3.00 per million audio tokens, and $1.00 per million image or video tokens. Google also gives approximate audio and visual rates of $0.005 per minute for audio input and $0.002 per minute for image or video input.
Output pricing is $4.50 per million text tokens and $12.00 per million audio tokens. The approximate generated-audio rate is $0.018 per minute. These per-minute figures are useful for estimating voice workloads, but actual billing depends on the provider’s token accounting and the amount of content processed.
A free tier is available subject to applicable usage limits. Search grounding includes 5,000 free search requests per month shared across Gemini 3.x models. Additional Search grounding requests are priced at $14 per 1,000 queries according to the supplied pricing research.
The cost trade-off is straightforward: audio output costs more per token than text output, but native audio can remove the need for a separate speech-generation stage and can reduce interaction latency. Teams should estimate both model usage and external tool or search activity when calculating the total cost of a voice agent.
Main strengths
- Low-latency voice interaction: The model is specifically designed for live dialogue and audio-to-audio applications.
- Native audio: It can receive and generate audio directly rather than relying entirely on a text-only intermediary.
- Broad multimodal input: An agent can combine spoken conversation with text, images, and video.
- Responsive tool use: Asynchronous function calling lets external operations run without automatically freezing the conversation.
- Current-information support: Search grounding is available for applications that need web-connected answers.
- Stable availability: The model is generally available rather than being limited to preview status.
- Large session capacity: The documented input and output limits support substantial conversational context and application instructions.
Main limitations
Gemini 3.8 Live’s specialization also defines its limits. It does not support structured outputs, so it is not a natural choice when every response must conform to a strict JSON schema. Applications can request text and use their own parsing or validation, but the model does not provide the documented structured-output capability.
Prompt caching and batch processing are not supported. This makes the model less appropriate for large offline workloads with repeated prefixes, scheduled processing, or cost-sensitive bulk inference. It also has no built-in code execution, file search, URL context, or Google Maps grounding in the supplied capability information.
Image generation is unavailable, and the model cannot be used as a single-model solution for a workflow that needs both live voice conversation and visual asset creation. Proactive audio is permanently enabled, which can complicate applications that require a fully manual push-to-talk interaction pattern.
Although the model supports interleaved reasoning, it is optimized for responsiveness rather than for long, deliberative research tasks. The supplied editorial assessment rates its reasoning at 7 out of 10 and coding at 6 out of 10, but these are comparative editorial scores, not Google-published benchmarks. Coding assistance is possible through text interaction and tool calls, but the model has no built-in code execution capability.
When to choose Gemini 3.8 Live
Choose Gemini 3.8 Live when the central requirement is a fast, natural conversation with streamed audio. Suitable examples include:
- Customer-service and support voice agents
- Voice-controlled applications
- Interactive tutoring and language practice
- Multimodal assistants that can discuss images or video during a conversation
- Field-support agents that combine spoken instructions with visual context
- Tool-using assistants that need to call business systems while maintaining dialogue
It is particularly attractive when a separate speech-recognition and text-to-speech pipeline would introduce unwanted delay or operational complexity. Search grounding can also help when the assistant must answer questions involving current external information.
When another option may be more appropriate
Use a different model or workflow when strict machine-readable output is the priority. Gemini 3.8 Live does not support structured outputs, so a model with native schema-constrained generation is a better fit for dependable JSON extraction or automated data pipelines.
A batch-oriented model is more appropriate for offline document processing, scheduled jobs, or high-volume inference. A model with prompt caching may be more economical when the same large instructions or context are sent repeatedly. Applications that need image generation should use a model designed to produce images rather than treating Gemini 3.8 Live as a general media-generation system.
For highly deliberative research or complex coding workflows, a model optimized for extended reasoning, code execution, or repository tools may offer a better capability trade-off even if it responds more slowly. Gemini 3.8 Live is the stronger choice when immediacy and spoken interaction outweigh those requirements.
Migration and design checklist
Teams moving from the older gemini-3.1-flash-live-preview model should change the model identifier to gemini-3.8-live and review session configuration rather than assuming complete behavioral compatibility.
- Remove unsupported
thinking_levelandthinking_configsettings. - Review asynchronous function-calling behavior and decide whether any tool must block the conversation.
- Design for permanently enabled proactive audio.
- Handle interruptions and partial updates using the Live API turn-completion rules.
- Enable output audio transcription when written transcripts are required.
- Do not depend on structured outputs, prompt caching, batch processing, or code execution.
- Estimate audio token costs separately from text, image, and video usage.
Overall, Gemini 3.8 Live is best understood as a stable real-time conversation engine. Its value comes from combining low-latency native audio, multimodal input, streaming sessions, search grounding, and responsive tool use. Its design is less suitable for strict schemas, offline throughput, visual generation, or workflows where extended reasoning and execution tools matter more than conversational speed.

