Gemini 3.5 Transcribe is Google's specialized speech-to-text model for turning prerecorded audio into written transcripts. Rather than serving as a general-purpose conversational or reasoning model, it focuses on accurately and efficiently extracting spoken language from audio files. It is available through the Gemini API under the canonical model ID gemini-3.5-transcribe and is listed as generally available.
The model is part of Google's Gemini 3.5 Audio family. Its role in that lineup is deliberately narrow: it handles batch or non-streaming transcription, while Google's separate gemini-3.5-transcribe-live endpoint is intended for live transcription. That distinction matters when selecting an implementation, because Gemini 3.5 Transcribe is not designed to maintain a real-time audio session.
What Gemini 3.5 Transcribe does
At its simplest, the model accepts an audio file and returns a text transcript. It can also add structure and metadata that make the result more useful than raw speech recognition. Supported capabilities include automatic language detection, multilingual transcription, multilingual code-switching, speaker diarization, word-level timestamps, smart formatting, and custom vocabulary biasing.
Speaker diarization identifies and labels different voices in a recording. This is useful for interviews, meetings, panels, and customer-service recordings. The model supports up to eight speakers, although attribution for three or more speakers is described as experimental. Word-level timestamps associate individual words with their position in the audio, which can help with subtitles, searchable recordings, editing workflows, and synchronized transcripts.
Smart transcription is intended to produce a cleaner reading transcript by removing disfluencies and applying formatting. This makes the output more suitable for notes, dictation, and document workflows than a strictly verbatim transcript. If preserving every hesitation or verbal filler is important, users should evaluate the resulting output carefully rather than assuming that smart formatting is a verbatim mode.
Supported inputs and outputs
Gemini 3.5 Transcribe accepts audio input and produces text output. It does not generate audio, images, video, music, embeddings, or other non-text outputs. The supplied specifications classify audio as its only multimodal input and text as its only output modality.
| Capability | Gemini 3.5 Transcribe |
|---|---|
| Primary task | Prerecorded, non-streaming audio transcription |
| Audio input | Supported |
| Text output | Supported |
| Image, video, or text input | Not listed as supported for this specialized model |
| Audio output | Not supported |
| Streaming | Not supported by this model |
| Tool or function calling | Not supported |
The model supports more than 85 languages and can handle multilingual code-switching, meaning that a recording can move between languages rather than remaining entirely in one language. The research identifies language coverage and code-switching as supported features, but does not provide a complete language-by-language list in the supplied specifications.
Audio limits and transcription options
A single request can process an audio file of up to one hour. Two optional features reduce that maximum: requests using speaker diarization or word-level timestamps are limited to 30 minutes. Applications processing longer recordings therefore need to divide them into segments when those annotations are required.
Custom vocabulary biasing can help the recognizer handle specialized terms, such as product names, medical terminology, proper nouns, or internal project language. The supplied documentation states that custom vocabulary supports up to 1,000 terms. Custom vocabulary cannot be combined with speaker diarization or word-level timestamps, so an application may need separate transcription passes if it requires both specialized vocabulary handling and detailed annotations.
The model has a listed context length of 96,000 tokens and a maximum output limit of 32,000 tokens. These are API model limits rather than a guarantee that every one-hour recording will fit into a single response, since the amount of generated text depends on speaking rate, language, formatting, and the requested annotations.
Pricing and value
Gemini 3.5 Transcribe has a free tier according to the supplied research. On the paid tier, the listed input price is $2.00 per one million audio input tokens, with an approximate supplied estimate of $0.003 per minute. Text output is listed at $12.00 per one million text output tokens, with an approximate supplied estimate of $0.002 per minute.
| Price component | Listed price |
|---|---|
| Audio input | $2.00 per 1 million audio input tokens |
| Text output | $12.00 per 1 million text output tokens |
| Free tier | Available |
Actual cost depends on token usage and the amount of transcript text returned. The per-minute figures are approximate estimates supplied with the model research, so production budgeting should use the token-based pricing and confirm current Google pricing before deployment. The model's editorial cost score is 8 out of 10, but that score is an evaluation in the supplied data, not a Google-published benchmark or price guarantee.
Strengths and trade-offs
The main strength of Gemini 3.5 Transcribe is specialization. A transcription-focused model is a more direct fit for turning recordings into text than a general-purpose model that happens to accept audio. The feature set covers several common production requirements in one service: multilingual recognition, code-switching, speaker labels, word timing, formatting, and domain vocabulary.
Its speed is another practical advantage. The supplied editorial evaluation gives the model a speed score of 9 out of 10 and a cost score of 8 out of 10. These scores should be treated as comparative editorial judgments rather than provider-published performance measurements. They indicate that the model is positioned for fast, economical transcription, not that Google guarantees a particular latency or accuracy level.
The trade-offs are equally important. Gemini 3.5 Transcribe is not a general reasoning model: its reasoning score is listed as 2 out of 10 and its coding score as 1 out of 10 in the supplied evaluation. It has no listed tool or function-calling support, no streaming capability, and no audio output. It is therefore a poor fit for an agent that must listen continuously, answer questions about audio in real time, call external tools, write substantial code, or speak a response.
When to choose Gemini 3.5 Transcribe
Choose Gemini 3.5 Transcribe when the central task is converting existing audio into usable text. It is particularly suitable for:
- Transcribing interviews, meetings, lectures, podcasts, and recorded calls.
- Creating transcripts across a large multilingual audience.
- Handling recordings that switch between languages.
- Producing speaker-labeled transcripts for conversations and panels.
- Generating word-timed text for captions, searchable media, or editing tools.
- Transcribing specialized material with up to 1,000 custom vocabulary terms, when diarization and word timestamps are not also required.
- Building batch transcription pipelines where live streaming is unnecessary.
It may not be the right choice when the application needs continuous, low-latency transcription. In that case, the separate Gemini 3.5 Transcribe Live model is the more relevant Google option according to the supplied research. A general-purpose multimodal model may be more appropriate when transcription is only one step in a larger workflow involving audio question answering, reasoning over the recording, tool calls, or code generation. Gemini 3.5 Transcribe should be selected for its transcription behavior, not treated as a replacement for those broader model categories.
Practical limitations to plan for
Recording length and optional annotations create the most significant operational constraint. Files up to one hour are supported in general, but diarization and word-level timestamps reduce the maximum to 30 minutes. A processing pipeline should segment longer recordings and decide whether it needs speaker labels, word timing, custom vocabulary, or a combination that the model does not allow in one request.
Speaker attribution also deserves quality control. Support for up to eight speakers is specified, but attribution for three or more speakers is experimental. For important legal, medical, or business records, applications should provide a review step rather than assuming that every label is correct.
Similarly, smart transcription is useful when readability is more important than strict verbatim fidelity, but the removal of disfluencies means that it should be matched to the intended use. A polished summary transcript and an evidentiary verbatim record have different requirements.
Overall assessment
Gemini 3.5 Transcribe is best understood as a focused audio-to-text service within Google's Gemini catalog. Its combination of multilingual support, code-switching, diarization, word-level timing, formatting, and vocabulary biasing gives it a broad set of transcription features while keeping the product's purpose clear. The one-hour general audio limit and 30-minute annotation limit are important implementation details, as are the incompatibilities between custom vocabulary and detailed annotations.
For prerecorded audio, the model offers a strong balance of feature coverage, speed, and stated cost efficiency. Its limitations are not missing general intelligence so much as deliberate specialization: it does not stream, call tools, reason broadly, generate speech, or serve as a general audio agent. Teams that need those capabilities should use a more suitable live or general-purpose model, while teams primarily needing reliable, structured transcripts can use Gemini 3.5 Transcribe as the focused option.

