Gemini 3.5 Audio

Gemini 3.5 Transcribe

by Google DeepMind · Generally available (GA)

Google's Gemini 3.5 Transcribe is a specialized, generally available model for prerecorded audio transcription. It supports more than 85 languages, multilingual code-switching, speaker diarization, word-level timestamps, smart formatting, and custom vocabulary biasing. Audio files can be up to one hour, although diarization and word-level timestamp requests are limited to 30 minutes. Paid pricing is listed at $2 per million audio input tokens and $12 per million text output tokens, with a free tier available.

Text Reasoning Coding
Gemini 3.5 Transcribe is a generally available Google model designed specifically for non-streaming audio transcription. It converts prerecorded audio into text, automatically detects languages, supports multilingual speech and speaker labeling, and can return optional word-level timing information. Audio files can be up to one hour per request, with a 30-minute limit when speaker diarization or word-level timestamps are enabled.
Outputs

What Gemini 3.5 Transcribe can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.5 Audio
Model type Other
Context window 96K tokens
Maximum output 32K tokens
Release date 2026-08-26
Status Generally available (GA)
Knowledge cutoff notes

Google's public model documentation and model card specify the context window and transcription capabilities but do not state a separate training-data knowledge cutoff for this specialized speech-to-text model.

Model notes

The canonical model ID is gemini-3.5-transcribe. It processes audio files through the Gemini API and returns text with optional word annotations. Audio files can be up to one hour per request, but requests using speaker diarization or word-level timestamps are limited to 30 minutes. Speaker diarization supports up to eight speakers, while attribution for three or more speakers is experimental. Custom vocabulary supports up to 1,000 terms and is incompatible with diarization and word-level timestamps. Smart transcription removes disfluencies and applies formatting. The separate gemini-3.5-transcribe-live endpoint is a distinct streaming model and should not be treated as the same exact model.

Cost

Model pricing

Input $2.00 per 1M audio input tokens, approximately $0.003 per minute; free tier available
Output $12.00 per 1M text output tokens, approximately $0.002 per minute; free tier available
Model guide

Gemini 3.5 Transcribe: Fast, Specialized Speech-to-Text for Multilingual Audio

Gemini 3.5 Transcribe is Google's dedicated, generally available speech-to-text model for converting prerecorded audio into readable transcripts. It supports more than 85 languages, multilingual code-switching, speaker diarization, word-level timestamps, smart formatting, and custom vocabulary biasing. The model accepts audio files of up to one hour, although diarization and word-level timestamp requests are limited to 30 minutes. It produces text rather than audio or other non-text outputs and is optimized for fast, cost-conscious transcription rather than general reasoning, coding, tool use, or real-time conversation.

Gemini 3.5 Transcribe is Google's specialized speech-to-text model for turning prerecorded audio into written transcripts. Rather than serving as a general-purpose conversational or reasoning model, it focuses on accurately and efficiently extracting spoken language from audio files. It is available through the Gemini API under the canonical model ID gemini-3.5-transcribe and is listed as generally available.

The model is part of Google's Gemini 3.5 Audio family. Its role in that lineup is deliberately narrow: it handles batch or non-streaming transcription, while Google's separate gemini-3.5-transcribe-live endpoint is intended for live transcription. That distinction matters when selecting an implementation, because Gemini 3.5 Transcribe is not designed to maintain a real-time audio session.

What Gemini 3.5 Transcribe does

At its simplest, the model accepts an audio file and returns a text transcript. It can also add structure and metadata that make the result more useful than raw speech recognition. Supported capabilities include automatic language detection, multilingual transcription, multilingual code-switching, speaker diarization, word-level timestamps, smart formatting, and custom vocabulary biasing.

Speaker diarization identifies and labels different voices in a recording. This is useful for interviews, meetings, panels, and customer-service recordings. The model supports up to eight speakers, although attribution for three or more speakers is described as experimental. Word-level timestamps associate individual words with their position in the audio, which can help with subtitles, searchable recordings, editing workflows, and synchronized transcripts.

Smart transcription is intended to produce a cleaner reading transcript by removing disfluencies and applying formatting. This makes the output more suitable for notes, dictation, and document workflows than a strictly verbatim transcript. If preserving every hesitation or verbal filler is important, users should evaluate the resulting output carefully rather than assuming that smart formatting is a verbatim mode.

Supported inputs and outputs

Gemini 3.5 Transcribe accepts audio input and produces text output. It does not generate audio, images, video, music, embeddings, or other non-text outputs. The supplied specifications classify audio as its only multimodal input and text as its only output modality.

CapabilityGemini 3.5 Transcribe
Primary taskPrerecorded, non-streaming audio transcription
Audio inputSupported
Text outputSupported
Image, video, or text inputNot listed as supported for this specialized model
Audio outputNot supported
StreamingNot supported by this model
Tool or function callingNot supported

The model supports more than 85 languages and can handle multilingual code-switching, meaning that a recording can move between languages rather than remaining entirely in one language. The research identifies language coverage and code-switching as supported features, but does not provide a complete language-by-language list in the supplied specifications.

Audio limits and transcription options

A single request can process an audio file of up to one hour. Two optional features reduce that maximum: requests using speaker diarization or word-level timestamps are limited to 30 minutes. Applications processing longer recordings therefore need to divide them into segments when those annotations are required.

Custom vocabulary biasing can help the recognizer handle specialized terms, such as product names, medical terminology, proper nouns, or internal project language. The supplied documentation states that custom vocabulary supports up to 1,000 terms. Custom vocabulary cannot be combined with speaker diarization or word-level timestamps, so an application may need separate transcription passes if it requires both specialized vocabulary handling and detailed annotations.

The model has a listed context length of 96,000 tokens and a maximum output limit of 32,000 tokens. These are API model limits rather than a guarantee that every one-hour recording will fit into a single response, since the amount of generated text depends on speaking rate, language, formatting, and the requested annotations.

Pricing and value

Gemini 3.5 Transcribe has a free tier according to the supplied research. On the paid tier, the listed input price is $2.00 per one million audio input tokens, with an approximate supplied estimate of $0.003 per minute. Text output is listed at $12.00 per one million text output tokens, with an approximate supplied estimate of $0.002 per minute.

Price componentListed price
Audio input$2.00 per 1 million audio input tokens
Text output$12.00 per 1 million text output tokens
Free tierAvailable

Actual cost depends on token usage and the amount of transcript text returned. The per-minute figures are approximate estimates supplied with the model research, so production budgeting should use the token-based pricing and confirm current Google pricing before deployment. The model's editorial cost score is 8 out of 10, but that score is an evaluation in the supplied data, not a Google-published benchmark or price guarantee.

Strengths and trade-offs

The main strength of Gemini 3.5 Transcribe is specialization. A transcription-focused model is a more direct fit for turning recordings into text than a general-purpose model that happens to accept audio. The feature set covers several common production requirements in one service: multilingual recognition, code-switching, speaker labels, word timing, formatting, and domain vocabulary.

Its speed is another practical advantage. The supplied editorial evaluation gives the model a speed score of 9 out of 10 and a cost score of 8 out of 10. These scores should be treated as comparative editorial judgments rather than provider-published performance measurements. They indicate that the model is positioned for fast, economical transcription, not that Google guarantees a particular latency or accuracy level.

The trade-offs are equally important. Gemini 3.5 Transcribe is not a general reasoning model: its reasoning score is listed as 2 out of 10 and its coding score as 1 out of 10 in the supplied evaluation. It has no listed tool or function-calling support, no streaming capability, and no audio output. It is therefore a poor fit for an agent that must listen continuously, answer questions about audio in real time, call external tools, write substantial code, or speak a response.

When to choose Gemini 3.5 Transcribe

Choose Gemini 3.5 Transcribe when the central task is converting existing audio into usable text. It is particularly suitable for:

  • Transcribing interviews, meetings, lectures, podcasts, and recorded calls.
  • Creating transcripts across a large multilingual audience.
  • Handling recordings that switch between languages.
  • Producing speaker-labeled transcripts for conversations and panels.
  • Generating word-timed text for captions, searchable media, or editing tools.
  • Transcribing specialized material with up to 1,000 custom vocabulary terms, when diarization and word timestamps are not also required.
  • Building batch transcription pipelines where live streaming is unnecessary.

It may not be the right choice when the application needs continuous, low-latency transcription. In that case, the separate Gemini 3.5 Transcribe Live model is the more relevant Google option according to the supplied research. A general-purpose multimodal model may be more appropriate when transcription is only one step in a larger workflow involving audio question answering, reasoning over the recording, tool calls, or code generation. Gemini 3.5 Transcribe should be selected for its transcription behavior, not treated as a replacement for those broader model categories.

Practical limitations to plan for

Recording length and optional annotations create the most significant operational constraint. Files up to one hour are supported in general, but diarization and word-level timestamps reduce the maximum to 30 minutes. A processing pipeline should segment longer recordings and decide whether it needs speaker labels, word timing, custom vocabulary, or a combination that the model does not allow in one request.

Speaker attribution also deserves quality control. Support for up to eight speakers is specified, but attribution for three or more speakers is experimental. For important legal, medical, or business records, applications should provide a review step rather than assuming that every label is correct.

Similarly, smart transcription is useful when readability is more important than strict verbatim fidelity, but the removal of disfluencies means that it should be matched to the intended use. A polished summary transcript and an evidentiary verbatim record have different requirements.

Overall assessment

Gemini 3.5 Transcribe is best understood as a focused audio-to-text service within Google's Gemini catalog. Its combination of multilingual support, code-switching, diarization, word-level timing, formatting, and vocabulary biasing gives it a broad set of transcription features while keeping the product's purpose clear. The one-hour general audio limit and 30-minute annotation limit are important implementation details, as are the incompatibilities between custom vocabulary and detailed annotations.

For prerecorded audio, the model offers a strong balance of feature coverage, speed, and stated cost efficiency. Its limitations are not missing general intelligence so much as deliberate specialization: it does not stream, call tools, reason broadly, generate speech, or serve as a general audio agent. Teams that need those capabilities should use a more suitable live or general-purpose model, while teams primarily needing reliable, structured transcripts can use Gemini 3.5 Transcribe as the focused option.


Answers to Frequently Asked Questions

Can custom vocabulary be used with speaker diarization or word-level timestamps?
No. Custom vocabulary supports up to 1,000 terms but cannot be combined with speaker diarization or word-level timestamps in the same request. Applications requiring both specialized vocabulary and detailed annotations may need to run separate transcription passes.
How many languages and speakers does Gemini 3.5 Transcribe support?
The model supports more than 85 languages and can transcribe recordings that switch between languages. It supports speaker diarization for up to eight speakers, although speaker attribution for three or more speakers is considered experimental and should be reviewed for important use cases.
Does Gemini 3.5 Transcribe support live or streaming transcription?
No. Gemini 3.5 Transcribe is designed for batch, non-streaming transcription of prerecorded audio. For continuous or real-time transcription, Google provides the separate gemini-3.5-transcribe-live endpoint.
What are the audio duration limits for Gemini 3.5 Transcribe?
A single request can process up to one hour of audio. When speaker diarization or word-level timestamps are enabled, the maximum duration is 30 minutes. Longer recordings must be divided into segments when these annotations are required.
What is Gemini 3.5 Transcribe used for?
Gemini 3.5 Transcribe is Google's specialized speech-to-text model for converting prerecorded audio into text. It supports multilingual transcription, language detection, code-switching, speaker diarization, word-level timestamps, smart formatting, and custom vocabulary biasing.


Sources 6
Provider

About Google DeepMind