What is Gemini 2.5 Flash Live?
Gemini 2.5 Flash Live is a preview model from Google DeepMind for building real-time conversational agents with the Gemini Live API. Its canonical model identifier is gemini-2.5-flash-native-audio-preview-12-2025. Google lists the model as released on December 12, 2025.
Unlike a conventional text model that receives a completed prompt and returns a completed answer, this model is designed for bidirectional streaming. An application can continuously send audio, video, or text while the model processes the session and produces audio or text responses. This makes it particularly relevant to voice assistants, interactive support agents, tutoring systems, and applications that need to react while a conversation is still taking place.
The model's native-audio design means that spoken input and spoken output can be handled within the Live API workflow without requiring an application to assemble separate speech-to-text and text-to-speech models for the core conversation. That does not remove the need for application-level session management, interruption handling, or audio transport, but it can simplify the central conversational pipeline.
Where it fits in Google's model lineup
Gemini 2.5 Flash Live occupies a specialized position in the Gemini 2.5 family. It is not simply the general gemini-2.5-flash endpoint with audio enabled. It is a separate native-audio preview model intended for Live API sessions.
This distinction matters when configuring an application. Developers should use the exact native-audio model identifier rather than assuming that the standard Gemini 2.5 Flash model exposes the same streaming audio behavior. The model is also distinct from the retired gemini-live-2.5-flash-preview identifier. Google has newer Live API models, so this preview remains most appropriate when its documented behavior and pricing match a particular application requirement.
Google currently limits access to Gemini 2.5 models for users who have actively used them in the past. Gemini 2.5 Flash Live has no announced shutdown date in the supplied documentation, but its preview status means availability, quotas, behavior, and API details can change.
Supported inputs and outputs
The model supports three input modalities:
- Audio: Continuous spoken or other audio input for voice interaction.
- Video: Visual input for video-aware conversations.
- Text: Text instructions, context, and conversational content.
It produces two output modalities:
- Audio: Generated spoken responses for direct voice interaction.
- Text: Text responses or accompanying transcripts that applications can display, store, or use for interface features.
Audio output is the model's defining capability. It is not an image-generation, video-generation, or music-generation model. The supplied specifications list image output, video output, music output, embedding output, and action output as unsupported.
Core capabilities for live applications
Gemini 2.5 Flash Live combines several capabilities inside a streaming conversation:
- Native audio generation: The model can return spoken responses rather than only text that must be converted by another service.
- Audio reasoning: It can interpret audio as part of the model interaction instead of treating a voice application as a purely text-based workflow.
- Thinking and reasoning: Google documents thinking support for the model, allowing it to reason before producing a response. The exact behavior and latency impact should be tested in the target Live API application.
- Function calling: The model can request application-defined tools, such as checking an account, retrieving a record, or controlling a supported workflow. The application remains responsible for executing the function and returning its result.
- Search grounding: The model can use search grounding when an application needs information beyond its training cutoff. Search grounding supplements the model with external information; it does not change the underlying cutoff.
- Bidirectional streaming: The Live API supports an ongoing session in which input and output events can move in both directions rather than waiting for a single completed request.
These capabilities make the model more suitable for an agent that must listen, reason, speak, and call a tool during one conversation than for a batch document-processing workflow.
Context, output, and technical limits
| Specification | Documented value |
|---|---|
| Model status | Preview |
| Context limit | 131,072 tokens |
| Maximum output | 8,192 tokens |
| Knowledge cutoff | January 2025 |
| Input | Audio, video, and text |
| Output | Audio and text |
| Structured outputs | Not supported |
| Prompt caching | Not supported |
| Batch API | Not supported |
The 131,072-token context limit provides room for substantial conversational context, instructions, and tool results, but it does not guarantee that every long session will be equally responsive or inexpensive. In a live application, developers should still manage session history and avoid retaining irrelevant content.
The 8,192-token output limit is a documented maximum, not a target response length. Real-time voice interfaces normally need concise turns, and longer reasoning or responses can affect perceived latency and cost. The model's January 2025 knowledge cutoff also means that current facts should be supplied through search grounding or another external source when accuracy depends on recent information.
Pricing and access
Google's documented pricing is separated by modality. Text input costs $0.50 per 1 million tokens. Audio and video input cost $3.00 per 1 million tokens. Text output costs $2.00 per 1 million tokens, while audio output costs $12.00 per 1 million tokens.
The pricing structure reflects an important practical trade-off: generated audio is substantially more expensive than generated text, and audio or video input is more expensive than text input. A voice application that continuously streams microphone audio and returns spoken answers can therefore consume budget faster than a text-only application, especially if it keeps sessions open unnecessarily or generates long responses.
Google lists free-tier availability as free of charge, subject to applicable limits and usage policies. Free access should not be treated as unlimited production capacity. Preview availability, quotas, account eligibility, and regional conditions may affect whether the model is practical for a deployed service.
Reasoning, coding, and tool use
Google documents reasoning or thinking support, which is useful when a live agent must interpret a complicated request, decide which function to call, or combine conversation with grounded information. However, this capability does not make Gemini 2.5 Flash Live a general-purpose autonomous system. The application must define tools, execute requested functions safely, validate results, and decide what information can be returned to the user.
The supplied editorial assessment gives the model a reasoning score of 7 out of 10 and a coding score of 6 out of 10. These are comparative editorial estimates, not Google-published benchmarks. The coding score suggests that coding-related conversation and tool workflows may be useful, but the model should not be selected primarily as a software-development model when a dedicated text or code workflow would be more appropriate.
Function calling is supported, but code execution is not documented as supported. The model also lacks file search, URL context, structured outputs, and batch processing. Consequently, a tool-using voice agent can be built around it, but the surrounding application must supply integrations for data access, code execution, file retrieval, validation, and durable business logic.
Speed and cost trade-offs
The supplied editorial ratings give Gemini 2.5 Flash Live a speed score of 9 out of 10 and a cost score of 7 out of 10. These ratings are subjective comparisons rather than provider benchmarks. They reflect the model's intended low-latency role, but actual responsiveness depends on network conditions, audio encoding, session design, tool latency, response length, and the amount of context retained.
Its strongest efficiency advantage is architectural: native audio can reduce the need to coordinate separate speech-recognition and speech-synthesis model calls for the main conversation. Its main cost disadvantage is that audio output is priced at $12.00 per million tokens, compared with $2.00 for text output. Applications that only need text should not use this model merely because it can also speak.
Best use cases
- Real-time voice assistants: Spoken interfaces that need natural turn-taking and immediate audio responses.
- Interactive customer service: Agents that listen to a caller, reason about the request, and call approved business functions.
- Voice tutoring and coaching: Conversational systems that provide spoken explanations, practice, or feedback.
- Video-aware assistants: Applications that need to discuss visual input while also maintaining a spoken conversation.
- Hands-free productivity tools: Voice-controlled workflows that connect speech with calendar, account, search, or other application functions.
- Search-grounded spoken agents: Assistants that need to answer with newer external information instead of relying only on the January 2025 model cutoff.
When another option may be more appropriate
Choose another model or architecture when the main requirement is not live audio interaction. A conventional text model is likely a better fit for ordinary text generation, structured data workflows, or applications that do not need spoken output. Gemini 2.5 Flash Live does not support structured outputs, so it is a poor choice when the application requires provider-enforced JSON conforming to a schema.
It is also unsuitable for image generation, batch processing, prompt-cached workloads, code execution, file search, or URL-context workflows according to the supplied model documentation. Applications needing these features should use a compatible model and supporting service rather than trying to add them to this endpoint.
A stable production system may also prefer a non-preview Live API model if one provides the required audio behavior and has more predictable availability. Conversely, this model can be a reasonable choice when native spoken responses, multimodal streaming, function calling, and the documented pricing are more important than preview stability or broad workflow features.
When to choose Gemini 2.5 Flash Live
Choose Gemini 2.5 Flash Live when the central product experience is a live, multimodal conversation: the user speaks or shares visual input, the model reasons during the session, and the application responds with speech. Its combination of audio input, audio output, Live API streaming, function calling, thinking support, and search grounding is more directly aligned with that use case than a text-only endpoint.
Before adopting it, verify access eligibility, preview terms, quotas, audio costs, and the behavior of interruption and tool-call events in a realistic prototype. For a text-first, highly structured, cache-heavy, batch, or code-execution workload, the model's specialized strengths do not compensate for its documented limitations.

