What is Gemini 3.5 Live Translate?
Gemini 3.5 Live Translate is Google's preview model for translating spoken conversations as they happen. It is designed for continuous audio streams rather than ordinary text prompts and turn-based chatbot exchanges. A caller, meeting participant, traveler, or customer can speak naturally, while the model generates translated speech in the selected target language.
The model's canonical identifier is gemini-3.5-live-translate-preview. Google provides it through the Gemini Live API and Google AI Studio. Its main role in Google's current model lineup is highly specialized: it focuses on fast speech translation instead of general-purpose question answering, long-form reasoning, coding, or tool-using workflows.
Google describes support for more than 70 languages and more than 2,000 language pairs in its model materials. The model automatically detects supported spoken languages and uses a target-language setting to control the translation direction. Translation is bidirectional, so it can support conversations in which participants speak different languages.
How the live translation model works
Gemini 3.5 Live Translate receives speech audio through a streaming connection and produces translated speech audio as its primary response. This audio-to-audio design avoids requiring an application to wait for a complete recording, transcribe it, translate the text, and then synthesize a separate response. The result is intended to feel closer to an interpreter working during a live conversation.
Applications can optionally request transcripts of both the incoming speech and the translated output. These transcripts are useful for captions, searchable meeting records, accessibility features, debugging, and interfaces that need both spoken and written translation.
- Input: live speech audio only in the Live Translation configuration.
- Primary output: translated speech audio.
- Optional output: text transcripts of the input and translated speech.
- Language handling: automatic detection of supported spoken languages and a configured target language.
- Conversation behavior: continuous stream processing rather than conventional request-and-response turns.
- Voice behavior: Google says the output is designed to preserve aspects of intonation, pacing, and pitch.
The model also includes an option controlling whether speech that is already in the target language should be echoed back. That setting can be useful in conversations where the application wants both sides of an exchange to remain audible, but Google notes that echoing can introduce artifacts in noisy environments.
Supported modalities and API boundaries
Although the model belongs to a multimodal model family, Gemini 3.5 Live Translate is deliberately restricted in this use case. It accepts audio input and returns audio output, with optional text transcription. It does not accept images, video, text prompts, or files for the translation workflow.
This narrow interface is important because the model is optimized for latency and translation consistency. It is not a general Gemini Live assistant that can be expanded into an agent with external actions. The supplied documentation says that the model does not support function calling, tools, Google Search grounding, Google Maps grounding, code execution, file search, URL context, structured outputs, thinking, batch processing, or context caching.
As a result, an application should not select this model when it needs a JSON response, a database lookup, web research, document processing, or an assistant that can perform actions. A separate orchestration layer or a different model would be needed for those tasks.
Context, output limits, and technical specifications
Google's model documentation lists an input token limit of 131,072 and a maximum output limit of 65,536 tokens. The Google DeepMind model card describes these figures in rounded terms as a context window of up to 128K tokens and output capacity of 64K tokens. These limits describe the model's documented token capacity; live translation itself is primarily experienced as a continuous audio stream rather than as a large text-generation request.
| Specification | Documented detail |
|---|---|
| Model ID | gemini-3.5-live-translate-preview |
| Provider | Google DeepMind |
| Status | Preview |
| Input | Speech audio |
| Primary output | Translated speech audio |
| Optional output | Input and output text transcripts |
| Language coverage | More than 70 languages, according to Google |
| Documented input limit | 131,072 tokens |
| Documented maximum output | 65,536 tokens |
| Knowledge cutoff | January 2025, according to the model card |
The model has no documented reasoning or thinking mode. That is not a missing convenience feature so much as a reflection of its purpose: adding deliberative reasoning would generally work against the low-latency behavior needed for live interpretation. It also has little relevance to coding. It cannot execute code, call tools, or provide a structured programming workflow, and should not be evaluated as a coding model.
Pricing and speed versus cost
Google lists standard paid pricing of $3.50 per 1 million input audio tokens and $21.00 per 1 million output audio tokens. Google estimates audio usage at 25 tokens per second. Using that estimate, input audio costs approximately $0.0053 per minute and output audio costs approximately $0.0315 per minute. If an application sends and receives audio continuously, the combined estimate is about $0.0368 per minute.
Output audio is substantially more expensive per token than input audio, so applications should account for both sides of a conversation rather than calculating only the cost of incoming speech. Actual usage will depend on how much audio is transmitted, how long participants speak, whether silence is streamed, and how much translated output is produced.
The main trade-off is specialization. A live audio model can be more suitable for an interactive conversation than a pipeline that separately performs speech recognition, text translation, and speech synthesis, because the single streaming workflow is designed around low latency. However, this convenience comes with a narrower capability set and audio pricing that may not be ideal for offline or high-volume document translation.
Strengths and limitations
Key strengths
- Real-time interaction: continuous stream processing is designed for conversations, calls, meetings, and other situations where waiting for a complete utterance is undesirable.
- Direct speech output: translated audio can be delivered without requiring an application to build a separate speech-synthesis stage.
- Broad stated coverage: Google documents more than 70 supported languages and more than 2,000 language pairs.
- Natural delivery: Google says the model aims to preserve aspects of intonation, pacing, and pitch.
- Optional transcripts: applications can expose written captions or retain text representations alongside the audio.
- Simple specialization: the restricted interface can be an advantage for products that need a translation component rather than a general-purpose AI agent.
Important limitations
Google notes that voice characteristics can become inconsistent, especially after long pauses or during rapid exchanges involving several speakers. Automatic language detection may be less reliable with non-native accents, closely related languages, rapid language switching, or difficult audio conditions. Background-noise filtering is available, but ambient sound can still affect translation quality.
The model's preview status is another practical limitation. Production teams should test their language pairs, microphones, speaker patterns, accents, latency requirements, and failure handling before relying on it for important conversations. A human interpreter or a separate verification process may still be appropriate for legal, medical, safety-critical, or otherwise high-consequence communication.
Its lack of tools, structured output, text input, file support, and batch processing also limits its scope. It is not suitable for translating a collection of documents, analyzing an uploaded recording as a batch job, or returning machine-readable translation objects with additional metadata.
Best use cases
Gemini 3.5 Live Translate is best suited to applications where spoken language must be translated with minimal delay:
- Multilingual calls and online meetings
- Travel, hospitality, and visitor-assistance tools
- Voice interpretation features in communication applications
- Live customer-support translation
- Streaming audio prototypes and conversational translation products
- Captioned conversations that need both translated audio and text
For example, a hospitality application could stream a guest's speech to the model, set the target language to the staff member's language, play the translated response, and display the transcript as a caption. A meeting application could use the same pattern to provide translated speech while retaining text for accessibility or review.
When to choose Gemini 3.5 Live Translate
Choose Gemini 3.5 Live Translate when the central requirement is low-latency, spoken translation and the application can work within an audio-only interface. It is particularly appropriate when natural voice output and ongoing conversation matter more than general reasoning, external actions, or extensive customization.
Another translation architecture may be more appropriate when the input is written text, recorded files, documents, or large batches. A general-purpose model is a better fit when users need explanations, summarization, image understanding, web search, coding, or tool use in the same interaction. A conventional speech-recognition, text-translation, and speech-synthesis pipeline may also be preferable when an organization needs tighter control over each stage, specialized terminology handling, or independent replacement of individual components.
In short, this model should be evaluated as a focused real-time interpreter, not as a general Gemini assistant. Its value comes from combining streaming speech input with translated speech output; its limitations come from that same specialization.

