Gemini 3.5 Audio

Gemini 3.5 Live Translate

by Google DeepMind · Preview

Google's Gemini 3.5 Live Translate is a preview Live API model for continuous, bidirectional speech-to-speech translation. It supports more than 70 languages, natural audio output, optional transcripts, and real-time streaming, but does not support tools, structured outputs, image input, video input, or web search.

Text Speech Reasoning Coding
Gemini 3.5 Live Translate is a specialized Gemini 3.5 model for continuous, real-time speech-to-speech translation. Rather than functioning as a general chatbot or reasoning model, it acts more like a streaming interpreter: it listens to spoken audio, detects supported languages, and returns translated speech with low latency. The preview model is available through the Gemini Live API and Google AI Studio.
Outputs

What Gemini 3.5 Live Translate can produce

Text Speech
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
10/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.5 Audio
Model type Multimodal
Context window 131K tokens
Maximum output 66K tokens
Knowledge cutoff January 2025
Release date 2026-06-09
Status Preview
Knowledge cutoff notes

The Google DeepMind model card states that the knowledge cutoff date for Gemini 3.5 Live Translate is January 2025.

Model notes

Canonical model ID is gemini-3.5-live-translate-preview. This is a preview audio-to-audio translation model available through the Gemini Live API and Google AI Studio. Google documents support for more than 70 languages. Live Translation uses continuous stream processing and supports translation configuration with a target language code and optional target-language echoing. The exact translation model does not support tools, instructions, function calling, search grounding, Maps grounding, code execution, file search, URL context, structured outputs, thinking, caching, batch API, Flex inference, or Priority inference. Google estimates billing using 25 audio tokens per second and gives an effective combined audio rate of approximately $0.0368 per minute. The model card identifies the knowledge cutoff as January 2025 and notes possible voice inconsistency, language-detection difficulties, and background-noise artifacts.

Cost

Model pricing

Input $3.50 per 1 million input audio tokens, approximately $0.0053 per minute
Output $21.00 per 1 million output audio tokens, approximately $0.0315 per minute
Model guide

Gemini 3.5 Live Translate: Real-Time Speech-to-Speech Translation

Gemini 3.5 Live Translate is Google's preview audio-to-audio model for low-latency, bidirectional speech translation across more than 70 languages. It accepts live speech and produces translated speech, with optional input and output transcripts.

What is Gemini 3.5 Live Translate?

Gemini 3.5 Live Translate is Google's preview model for translating spoken conversations as they happen. It is designed for continuous audio streams rather than ordinary text prompts and turn-based chatbot exchanges. A caller, meeting participant, traveler, or customer can speak naturally, while the model generates translated speech in the selected target language.

The model's canonical identifier is gemini-3.5-live-translate-preview. Google provides it through the Gemini Live API and Google AI Studio. Its main role in Google's current model lineup is highly specialized: it focuses on fast speech translation instead of general-purpose question answering, long-form reasoning, coding, or tool-using workflows.

Google describes support for more than 70 languages and more than 2,000 language pairs in its model materials. The model automatically detects supported spoken languages and uses a target-language setting to control the translation direction. Translation is bidirectional, so it can support conversations in which participants speak different languages.

How the live translation model works

Gemini 3.5 Live Translate receives speech audio through a streaming connection and produces translated speech audio as its primary response. This audio-to-audio design avoids requiring an application to wait for a complete recording, transcribe it, translate the text, and then synthesize a separate response. The result is intended to feel closer to an interpreter working during a live conversation.

Applications can optionally request transcripts of both the incoming speech and the translated output. These transcripts are useful for captions, searchable meeting records, accessibility features, debugging, and interfaces that need both spoken and written translation.

  • Input: live speech audio only in the Live Translation configuration.
  • Primary output: translated speech audio.
  • Optional output: text transcripts of the input and translated speech.
  • Language handling: automatic detection of supported spoken languages and a configured target language.
  • Conversation behavior: continuous stream processing rather than conventional request-and-response turns.
  • Voice behavior: Google says the output is designed to preserve aspects of intonation, pacing, and pitch.

The model also includes an option controlling whether speech that is already in the target language should be echoed back. That setting can be useful in conversations where the application wants both sides of an exchange to remain audible, but Google notes that echoing can introduce artifacts in noisy environments.

Supported modalities and API boundaries

Although the model belongs to a multimodal model family, Gemini 3.5 Live Translate is deliberately restricted in this use case. It accepts audio input and returns audio output, with optional text transcription. It does not accept images, video, text prompts, or files for the translation workflow.

This narrow interface is important because the model is optimized for latency and translation consistency. It is not a general Gemini Live assistant that can be expanded into an agent with external actions. The supplied documentation says that the model does not support function calling, tools, Google Search grounding, Google Maps grounding, code execution, file search, URL context, structured outputs, thinking, batch processing, or context caching.

As a result, an application should not select this model when it needs a JSON response, a database lookup, web research, document processing, or an assistant that can perform actions. A separate orchestration layer or a different model would be needed for those tasks.

Context, output limits, and technical specifications

Google's model documentation lists an input token limit of 131,072 and a maximum output limit of 65,536 tokens. The Google DeepMind model card describes these figures in rounded terms as a context window of up to 128K tokens and output capacity of 64K tokens. These limits describe the model's documented token capacity; live translation itself is primarily experienced as a continuous audio stream rather than as a large text-generation request.

SpecificationDocumented detail
Model IDgemini-3.5-live-translate-preview
ProviderGoogle DeepMind
StatusPreview
InputSpeech audio
Primary outputTranslated speech audio
Optional outputInput and output text transcripts
Language coverageMore than 70 languages, according to Google
Documented input limit131,072 tokens
Documented maximum output65,536 tokens
Knowledge cutoffJanuary 2025, according to the model card

The model has no documented reasoning or thinking mode. That is not a missing convenience feature so much as a reflection of its purpose: adding deliberative reasoning would generally work against the low-latency behavior needed for live interpretation. It also has little relevance to coding. It cannot execute code, call tools, or provide a structured programming workflow, and should not be evaluated as a coding model.

Pricing and speed versus cost

Google lists standard paid pricing of $3.50 per 1 million input audio tokens and $21.00 per 1 million output audio tokens. Google estimates audio usage at 25 tokens per second. Using that estimate, input audio costs approximately $0.0053 per minute and output audio costs approximately $0.0315 per minute. If an application sends and receives audio continuously, the combined estimate is about $0.0368 per minute.

Output audio is substantially more expensive per token than input audio, so applications should account for both sides of a conversation rather than calculating only the cost of incoming speech. Actual usage will depend on how much audio is transmitted, how long participants speak, whether silence is streamed, and how much translated output is produced.

The main trade-off is specialization. A live audio model can be more suitable for an interactive conversation than a pipeline that separately performs speech recognition, text translation, and speech synthesis, because the single streaming workflow is designed around low latency. However, this convenience comes with a narrower capability set and audio pricing that may not be ideal for offline or high-volume document translation.

Strengths and limitations

Key strengths

  • Real-time interaction: continuous stream processing is designed for conversations, calls, meetings, and other situations where waiting for a complete utterance is undesirable.
  • Direct speech output: translated audio can be delivered without requiring an application to build a separate speech-synthesis stage.
  • Broad stated coverage: Google documents more than 70 supported languages and more than 2,000 language pairs.
  • Natural delivery: Google says the model aims to preserve aspects of intonation, pacing, and pitch.
  • Optional transcripts: applications can expose written captions or retain text representations alongside the audio.
  • Simple specialization: the restricted interface can be an advantage for products that need a translation component rather than a general-purpose AI agent.

Important limitations

Google notes that voice characteristics can become inconsistent, especially after long pauses or during rapid exchanges involving several speakers. Automatic language detection may be less reliable with non-native accents, closely related languages, rapid language switching, or difficult audio conditions. Background-noise filtering is available, but ambient sound can still affect translation quality.

The model's preview status is another practical limitation. Production teams should test their language pairs, microphones, speaker patterns, accents, latency requirements, and failure handling before relying on it for important conversations. A human interpreter or a separate verification process may still be appropriate for legal, medical, safety-critical, or otherwise high-consequence communication.

Its lack of tools, structured output, text input, file support, and batch processing also limits its scope. It is not suitable for translating a collection of documents, analyzing an uploaded recording as a batch job, or returning machine-readable translation objects with additional metadata.

Best use cases

Gemini 3.5 Live Translate is best suited to applications where spoken language must be translated with minimal delay:

  • Multilingual calls and online meetings
  • Travel, hospitality, and visitor-assistance tools
  • Voice interpretation features in communication applications
  • Live customer-support translation
  • Streaming audio prototypes and conversational translation products
  • Captioned conversations that need both translated audio and text

For example, a hospitality application could stream a guest's speech to the model, set the target language to the staff member's language, play the translated response, and display the transcript as a caption. A meeting application could use the same pattern to provide translated speech while retaining text for accessibility or review.

When to choose Gemini 3.5 Live Translate

Choose Gemini 3.5 Live Translate when the central requirement is low-latency, spoken translation and the application can work within an audio-only interface. It is particularly appropriate when natural voice output and ongoing conversation matter more than general reasoning, external actions, or extensive customization.

Another translation architecture may be more appropriate when the input is written text, recorded files, documents, or large batches. A general-purpose model is a better fit when users need explanations, summarization, image understanding, web search, coding, or tool use in the same interaction. A conventional speech-recognition, text-translation, and speech-synthesis pipeline may also be preferable when an organization needs tighter control over each stage, specialized terminology handling, or independent replacement of individual components.

In short, this model should be evaluated as a focused real-time interpreter, not as a general Gemini assistant. Its value comes from combining streaming speech input with translated speech output; its limitations come from that same specialization.


Answers to Frequently Asked Questions

What are the best use cases and limitations of Gemini 3.5 Live Translate?
The model is designed for low-latency spoken translation in multilingual calls, meetings, travel and hospitality tools, customer support, communication apps, and captioned conversations. It is not a general-purpose assistant and does not support tools, function calling, web or Maps grounding, code execution, structured outputs, batch processing, file translation, or document workflows. Because it is a preview model, applications should test language pairs, accents, background noise, latency, and speaker patterns before using it for important conversations.
How much does Gemini 3.5 Live Translate cost?
Google lists pricing of $3.50 per 1 million input audio tokens and $21.00 per 1 million output audio tokens. Based on Google's estimate of 25 audio tokens per second, input audio costs approximately $0.0053 per minute, output audio costs about $0.0315 per minute, and continuously sending and receiving audio costs roughly $0.0368 per minute.
What is Gemini 3.5 Live Translate?
Gemini 3.5 Live Translate is Google's preview model for translating spoken conversations in real time. Its canonical model ID is gemini-3.5-live-translate-preview, and it is available through the Gemini Live API and Google AI Studio.
What input and output modalities does Gemini 3.5 Live Translate support?
The model accepts live speech audio and produces translated speech audio. Applications can also request transcripts of the incoming speech and translated output. The translation workflow does not support text prompts, images, video, files, or other input modalities.


Sources 5
Provider

About Google DeepMind