Gemini 2.5

Gemini 2.5 Pro TTS

by Google DeepMind · Preview; currently listed and accessible through the Gemini API, with restricted access to Gemini 2.5 models

Gemini 2.5 Pro TTS is a preview Google model specialized in converting text into expressive audio. It supports single-speaker and multi-speaker speech, style and pacing control, batch processing, and long-form narration, but does not provide general reasoning, tool use, multimodal input, or real-time Live API interaction.

Speech Reasoning Coding
Gemini 2.5 Pro TTS is a specialized Google DeepMind model for turning written scripts into natural-sounding speech. It is designed for quality and control rather than general-purpose reasoning or real-time conversation, with support for single-speaker and multi-speaker audio generation, style guidance, pacing, and long-form narration workflows.
Outputs

What Gemini 2.5 Pro TTS can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Batch API Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
6/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family Gemini 2.5
Model type Other
Context window 8K tokens
Maximum output 16K tokens
Release date 2025-05-20
Status Preview; currently listed and accessible through the Gemini API, with restricted access to Gemini 2.5 models
Knowledge cutoff notes

Google's model documentation for Gemini 2.5 Pro TTS does not publish a separate knowledge-cutoff date. The model is a specialized text-to-speech endpoint rather than a general knowledge and reasoning model.

Model notes

Canonical model ID is gemini-2.5-pro-preview-tts. The model accepts text-only input and produces audio-only output. It supports controllable single-speaker and multi-speaker speech generation, including guidance for style, tone, accent, and pacing. Google documents audio generation as supported, while caching, function calling, code execution, search grounding, structured outputs, Live API, and reasoning are not supported. The model is a preview endpoint and newer Gemini TTS models are available for new projects. Pricing uses separate text-input and audio-output token units. Editorial scores are comparative estimates for this specialized TTS model, not vendor benchmarks.

Cost

Model pricing

Input $1.00 per 1M text tokens standard; $0.50 per 1M text tokens batch
Output $20.00 per 1M audio tokens standard; $10.00 per 1M audio tokens batch
Model guide

Gemini 2.5 Pro TTS: High-Fidelity Speech for Long-Form Narration

Gemini 2.5 Pro TTS is Google's preview text-to-speech model for generating expressive, high-fidelity audio from text. It supports single-speaker and multi-speaker narration, controllable delivery, and structured workflows such as audiobooks, podcasts, voiceovers, and scripted customer-support content.

What is Gemini 2.5 Pro TTS?

Gemini 2.5 Pro TTS is Google's preview text-to-speech model for converting written text into generated speech. Its canonical Gemini API model identifier is gemini-2.5-pro-preview-tts. The model belongs to the Gemini 2.5 family, but it should not be confused with the general-purpose Gemini 2.5 Pro model: this endpoint is specialized for audio generation rather than broad reasoning, coding, or multimodal analysis.

The model is provided by Google and is available through the Gemini API. Google positions it for high-fidelity, controllable speech production, including long-form narration, podcasts, audiobooks, professional voiceovers, scripted customer-support responses, and other planned audio workflows.

Its preview status is important. Gemini 2.5 Pro TTS remains listed in Google's model documentation, but Google has also released newer Gemini TTS models. Teams starting a new production integration should therefore compare this model with the provider's current TTS offerings and review migration guidance before making a long-term commitment.

Modalities and speech capabilities

Gemini 2.5 Pro TTS accepts text input and produces audio output. It does not provide a general text response alongside the audio, and it is not a speech-to-speech model. The input and output design is deliberately narrow, which makes the model easier to evaluate for scripted voice production but unsuitable for applications that need image, audio, or video understanding.

  • Input: Text only.
  • Output: Generated speech audio only.
  • Speaker modes: Single-speaker and multi-speaker generation.
  • Speech control: Prompts and speech-generation settings can guide delivery characteristics such as tone, accent, speaking style, and pace.
  • General reasoning: Not supported as a model capability.
  • Audio generation: Supported.

Single-speaker generation is appropriate for narration, explainers, lessons, announcements, and voiceovers. Multi-speaker generation is useful for scripted conversations, interviews, character dialogue, and podcast-style formats. In both cases, the model is best treated as a renderer for prepared or structured text rather than as an autonomous conversation partner.

Where it fits in the Gemini lineup

The “Pro” name can suggest a general-purpose premium reasoning model, but Gemini 2.5 Pro TTS serves a different role. It is a dedicated speech-generation endpoint in the Gemini 2.5 family. Its value comes from converting a script into expressive audio, not from answering questions, writing code, searching the web, or operating tools.

This specialization creates a clear trade-off. A general Gemini model may be more suitable when an application must interpret documents, reason over information, generate a script, or coordinate a broader workflow. Gemini 2.5 Pro TTS becomes relevant after the text to be spoken is available and the primary requirement is polished, controllable speech.

Because the model is a preview endpoint and newer Gemini TTS options are available, its place in the catalog may change. Developers should verify the live model catalog, access conditions, and migration documentation when planning a new deployment.

Strengths and practical use cases

The main strength of Gemini 2.5 Pro TTS is its focus on expressive, high-quality narration. It is designed for cases where vocal clarity, natural prosody, and control over delivery matter more than the lowest possible latency.

  • Audiobooks: Generate extended narration from prepared chapters or sections.
  • Podcasts: Produce scripted introductions, narration, interviews, or multi-person dialogue.
  • Professional voiceovers: Create audio for marketing videos, educational material, product demonstrations, and presentations.
  • Interactive fiction: Render character dialogue with multiple speakers and directed delivery.
  • Customer-support content: Convert approved, scripted responses into spoken output.
  • Learning content: Create narrated lessons or explanations from structured text.

Multi-speaker support is especially useful when a project would otherwise need to generate and assemble separate voices manually. The ability to describe style, tone, accent, and pacing also gives developers a way to specify how a script should sound, rather than relying only on the words in the input.

These are suitability assessments based on the model's documented speech-generation scope. They should not be read as independent benchmark results or a guarantee that every voice, accent, or long-form production will meet a particular quality standard.

Limits, context, and API support

Google's documentation lists an input limit of 8,192 tokens and a maximum output limit of 16,384 tokens. Tokens are units used to measure processed text and generated content; they do not correspond exactly to words or characters. For long scripts, developers may need to divide content into sections and manage continuity between generated segments.

The model supports batch API processing, which is relevant for planned workloads such as audiobook chapters, collections of voiceovers, or large sets of scripted responses. It is not documented as a Live API model and should not be treated as a real-time, bidirectional voice system.

According to the supplied model documentation, Gemini 2.5 Pro TTS does not support function calling, code execution, search grounding, prompt caching, structured outputs, or Live API sessions. It also does not accept image, audio, or video input, and it does not natively produce text, images, video, embeddings, or structured JSON output.

These limitations mean that surrounding application logic must handle tasks such as script creation, factual research, validation, tool use, file processing, and audio post-production. A common architecture may use a general-purpose model or ordinary application code to prepare and validate a script, then send the final text to Gemini 2.5 Pro TTS for speech generation.

Pricing and cost trade-offs

Google's listed standard pricing is $1.00 per 1 million text input tokens and $20.00 per 1 million audio output tokens. Batch pricing is listed at $0.50 per 1 million text input tokens and $10.00 per 1 million audio output tokens.

Usage modeText inputAudio output
Standard$1.00 per 1 million text tokens$20.00 per 1 million audio tokens
Batch$0.50 per 1 million text tokens$10.00 per 1 million audio tokens

Input and output use different token units, so the figures should not be treated as a single combined rate. The output price is the more significant cost consideration for large audio workloads. Batch processing can reduce listed token rates when the application can tolerate asynchronous processing and does not require immediate responses.

There is no supplied evidence that the model is optimized for the lowest-latency or lowest-cost speech generation. Editorially, it is best viewed as a quality-focused option: its potential advantages are fidelity, controllability, and multi-speaker support, while a simpler or newer TTS endpoint may be preferable when throughput, current support, or predictable production lifecycle matters more.

Reasoning, coding, and tool support

Gemini 2.5 Pro TTS is not a reasoning model in the practical sense used for question answering or planning. It does not provide general-purpose text output, and its role is to generate speech from supplied text. Its editorial reasoning and coding assessments are therefore very low because those capabilities are outside the model's intended function, not because the model is being evaluated as a poor general language model.

The model does not support code execution, function calling, search grounding, or other documented tool-use features. It cannot independently retrieve information, call a business system, browse the web, or decide which external action to take. If an application needs those functions, they must be performed before speech generation or by another model and then passed into the TTS workflow as approved text.

When to choose Gemini 2.5 Pro TTS

Choose Gemini 2.5 Pro TTS when the central requirement is expressive audio from a prepared script and the project benefits from single-speaker or multi-speaker generation. It is a reasonable fit for planned narration, podcast production, audiobooks, educational voiceovers, and other workflows where audio quality and delivery control are more important than interactive response speed.

Consider another option when you need:

  • Real-time, bidirectional voice conversation or speech-to-speech interaction.
  • Audio, image, or video understanding as input.
  • General reasoning, coding, web research, or tool-using agents.
  • Native structured JSON or text output from the same model endpoint.
  • Function calling, code execution, search grounding, or Live API sessions.
  • A current non-preview TTS model with a clearer long-term migration path.

Compared with a general-purpose language model, Gemini 2.5 Pro TTS offers a narrower but more relevant capability when the final product is spoken audio. Compared with a lightweight or latency-oriented speech service, its documented positioning favors fidelity and expressive control rather than simply minimizing processing time or cost. The right choice depends on whether the project prioritizes narration quality, multi-speaker control, throughput, real-time interaction, or long-term model availability.

Overall assessment

Gemini 2.5 Pro TTS is best understood as a specialized, preview speech-generation model rather than a conversational Gemini assistant. It accepts text, produces audio, supports single-speaker and multi-speaker workflows, and provides controls for delivery style and pacing. Its 8,192-token input limit, 16,384-token output limit, batch support, and documented pricing make it suitable for structured production pipelines, while its lack of tools, reasoning, multimodal input, and real-time Live API support defines its boundaries.

For a project that already has reliable scripts and needs polished narration, it can be a useful choice. For applications that still need to research, reason, converse, or act before speaking, it should be treated as the audio-generation component of a larger system rather than as the complete solution.


Answers to Frequently Asked Questions

How much does Gemini 2.5 Pro TTS cost?
Google's listed standard pricing is $1.00 per 1 million text input tokens and $20.00 per 1 million audio output tokens. Batch pricing is $0.50 per 1 million text input tokens and $10.00 per 1 million audio output tokens. Audio output tokens are the more significant cost factor for large speech-generation workloads.
What are the context limits and API limitations of Gemini 2.5 Pro TTS?
Gemini 2.5 Pro TTS has a documented input limit of 8,192 tokens and a maximum output limit of 16,384 tokens. It supports batch processing but is not documented as a Live API model. It does not support image, audio, or video input, function calling, code execution, search grounding, prompt caching, structured outputs, or real-time bidirectional voice sessions.
What types of speech generation does Gemini 2.5 Pro TTS support?
The model accepts text input and produces audio output only. It supports both single-speaker and multi-speaker generation, with prompts and speech-generation settings that can guide tone, accent, speaking style, and pace. It is suitable for narration, audiobooks, podcasts, voiceovers, scripted customer-support content, and character dialogue.
What is Gemini 2.5 Pro TTS and what is its model identifier?
Gemini 2.5 Pro TTS is Google's preview text-to-speech model for converting written text into generated speech. Its canonical Gemini API model identifier is gemini-2.5-pro-preview-tts. It is specialized for audio generation rather than general reasoning, coding, or multimodal analysis.


Sources 6
Provider

About Google DeepMind