Gemini 2.5

Gemini 2.5 Flash TTS

by Google DeepMind · Preview; currently available with limited access conditions

Gemini 2.5 Flash TTS is a preview Google Gemini API model that converts text into controllable spoken audio. It supports style, pacing, accent, single- and multi-speaker synthesis, 8,192 input tokens, 16,384 output tokens, standard pricing of $0.50 per 1 million input tokens and $10 per 1 million output audio tokens, plus lower batch rates.

Speech Reasoning Coding
Gemini 2.5 Flash TTS is a specialized Google model for developers who need text converted into spoken audio with control over delivery. It accepts text input and produces audio output, with support for characteristics such as style, tone, accent, pacing, and speaker configuration. Its main appeal is the balance between generation speed, cost, and controllability, while its preview status and narrow text-to-speech focus limit where it is appropriate.
Outputs

What Gemini 2.5 Flash TTS can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Batch API Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Gemini 2.5
Model type Other
Context window 8K tokens
Maximum output 16K tokens
Release date 2025-05-20
Status Preview; currently available with limited access conditions
Model notes

The canonical model identifier is gemini-2.5-flash-preview-tts. Google describes the model as a fast, controllable TTS model for low-latency and cost-efficient applications. It accepts text-only input and produces audio-only output. Google documents single- and multi-speaker synthesis for the 2.5 TTS preview models. The current model page lists an 8,192-token input limit and 16,384-token output limit. Google states that Gemini 2.5 models are not deprecated but access may be limited to users who have actively used them; newer models are recommended for new projects. Editorial capability scores reflect the model's specialized TTS role and should not be compared directly with general-purpose language models without considering modality and task.

Cost

Model pricing

Input $0.50 per 1M input text tokens; batch: $0.25 per 1M input text tokens
Output $10.00 per 1M output audio tokens; batch: $5.00 per 1M output audio tokens
Model guide

Gemini 2.5 Flash TTS: Fast, Controllable Speech Generation for Voice Applications

Gemini 2.5 Flash TTS is Google's preview text-to-speech model for generating controllable spoken audio from text. It is designed for low-latency, cost-efficient voice assistants, narration, read-aloud features, and single- or multi-speaker dialogue rather than general-purpose reasoning or multimodal understanding.

What is Gemini 2.5 Flash TTS?

Gemini 2.5 Flash TTS is Google's preview text-to-speech model for turning written text into spoken audio. Its canonical Gemini API model identifier is gemini-2.5-flash-preview-tts. The model was released on May 20, 2025, and is positioned within Google's Gemini 2.5 audio model offerings.

Unlike a general-purpose Gemini model, it is not intended to answer questions, analyze images, write code, or conduct broad multimodal reasoning. Its job is narrower: generate speech that follows instructions about how the words should be delivered. A developer can use it for a voice interface, audiobook-style narration, read-aloud functionality, simulated interviews, or dialogue involving more than one speaker.

Google describes the model as fast, controllable, and suitable for low-latency and cost-efficient applications. Those are provider positioning claims; the practical choice should also account for the model's preview status, access conditions, quotas, and the requirements of the application.

How it fits into Google's model lineup

Gemini 2.5 Flash TTS is a specialized member of the Gemini 2.5 family rather than a general chat model. Its focus is speech synthesis, while Gemini Live models are designed for interactive, bidirectional audio conversations. The distinction matters: Flash TTS generates spoken audio from supplied text, whereas a Live API model is the more natural category for speech-to-speech interaction and ongoing conversational audio.

Google's documentation states that Gemini 2.5 models are not deprecated and will continue to be served until further notice, but it recommends newer model generations for new projects. Access to this preview model may also be limited to users who have previously used Gemini 2.5 models. Teams considering it for a new production system should therefore confirm current availability and migration guidance before committing to an integration.

Input, output, and speech capabilities

The model has a simple modality profile:

  • Input: text only.
  • Output: generated speech audio only.
  • Speaker configurations: single-speaker and multi-speaker synthesis are documented for the Gemini 2.5 TTS preview models.
  • Text output: not provided as a native model output.

Controllability is the model's main functional distinction. Prompts can guide delivery characteristics such as tone, style, accent, and pacing. For example, the same written script could be requested as a calm product explanation, a faster conversational response, or a multi-character exchange with different speakers. The model is therefore more useful than a basic fixed-voice converter when the application needs speech that reflects context or presentation style.

It should not be confused with a model that understands audio. Gemini 2.5 Flash TTS does not accept audio, images, or video input, so it cannot directly transcribe a recording, interpret a spoken request, or use visual context to determine how a script should be read.

Technical specifications

SpecificationVerified detail
Model IDgemini-2.5-flash-preview-tts
Release dateMay 20, 2025
StatusPreview
Input context limit8,192 tokens
Maximum output limit16,384 tokens
Input modalityText
Output modalityAudio
Function callingNot supported
Structured outputsNot supported
Search groundingNot supported
Live APINot supported
CachingNot supported
Batch APISupported according to the current model documentation

The token limits describe the model's text processing and generated-output capacity; they should not be interpreted as a promise about a particular duration of finished audio. Actual speech length depends on the supplied text, speaking style, language or accent behavior, and generation details.

Pricing and cost trade-offs

Google's current Gemini API pricing lists standard paid usage at $0.50 per 1 million input text tokens and $10.00 per 1 million output audio tokens. Batch pricing is listed at $0.25 per 1 million input text tokens and $5.00 per 1 million output audio tokens.

The standard rate is the relevant option for interactive applications that need prompt-to-speech responses with low latency. Batch pricing is potentially better suited to asynchronous workloads such as preparing a large collection of narration files, provided the application's timing requirements allow batch processing. Free-tier availability and eligibility are controlled by Google and may not match paid API access.

These prices are token-based rather than a simple per-minute audio rate. As a result, teams should estimate costs using representative scripts and output behavior instead of assuming that every minute of speech has the same price. The model's cost advantage is most relevant when compared with more capable or higher-fidelity speech options, but the supplied pricing does not establish a direct quality comparison.

Speed, reasoning, and coding capability

Gemini 2.5 Flash TTS is intended for fast speech generation, especially in applications where users are waiting for an audio response. Its specialization can make it a sensible choice when the text to be spoken has already been prepared by an application or another model. It is not, however, a reasoning engine for deciding what to say.

The supplied editorial evaluation rates its reasoning capability as 2 out of 10 and coding capability as 1 out of 10. These are editorial scores, not Google-published benchmarks. They reflect the model's specialized TTS role: it may follow instructions about delivery, but it should not be selected for complex analysis, software generation, tool orchestration, or general-purpose problem solving.

There is no function calling, search grounding, structured output, or Live API support. If an application needs current information, calculations, database access, external actions, or a structured decision before speech is generated, that work must be handled elsewhere in the system. A separate reasoning or application layer can prepare the text, after which Gemini 2.5 Flash TTS can synthesize the result.

Best use cases

The model is most suitable when text-to-speech is the central requirement and the application benefits from responsive generation or controllable delivery. Suitable examples include:

  • Voice assistants: convert already-prepared responses into spoken replies without requiring a general audio conversation model.
  • Read-aloud features: generate spoken versions of articles, help content, notifications, or educational material.
  • Narration: produce voice tracks for high-volume content workflows where prompt-controlled style and pacing are useful.
  • Podcasts and simulated interviews: create multi-character dialogue with different speaker configurations.
  • Localized content: generate speech for applications that need control over accent, delivery, or presentation style.
  • Batch audio production: create many narration assets asynchronously when immediate response time is not essential.

For a production workflow, developers should test representative scripts rather than evaluating only short demonstrations. Long passages, dialogue turns, unusual names, specialized terminology, and tightly controlled pacing can expose differences between a prototype and a dependable user experience.

Limitations and when to choose another option

The most important limitation is scope. Gemini 2.5 Flash TTS is not a general-purpose multimodal model and does not accept image, audio, or video input. It does not generate text as its native output, call functions, use search grounding, return structured outputs, or connect through the Live API. It is therefore a poor fit for an application that must listen to a user, interpret visual content, retrieve information, call tools, and respond conversationally in one model interaction.

A speech-to-speech or Live API option is more appropriate for genuinely bidirectional audio conversations. A general-purpose language model is more appropriate when the difficult part of the task is reasoning, coding, research, or generating the script itself. A newer recommended Gemini TTS model may be preferable for a new project if it offers a better support horizon, capabilities, or migration path, although the supplied research does not provide a detailed feature or price comparison with those newer models.

The preview label also introduces operational risk. Availability, quotas, supported access paths, and future migration requirements can change. Confirm those details with Google's current documentation before using the model as the sole speech engine for a critical product.

Bottom line

Gemini 2.5 Flash TTS is a focused option for turning text into expressive, controllable speech at a relatively low token-based cost. Its strengths are low-latency positioning, style and pacing control, single- and multi-speaker synthesis, and support for both standard and batch workflows. Its weaknesses are equally clear: preview availability, text-only input, audio-only output, no tool or Live API support, and very limited usefulness for reasoning or coding.

Choose it when an application already has the text and needs that text spoken quickly and flexibly. Choose another model type when the system must understand audio or visual inputs, reason over information, use external tools, or sustain a full two-way voice conversation.


Answers to Frequently Asked Questions

When should developers choose another model instead of Gemini 2.5 Flash TTS?
Choose another model when the application needs speech-to-speech interaction, audio or visual understanding, complex reasoning, coding, search, function calling, structured outputs, or a persistent two-way voice conversation. Gemini 2.5 Flash TTS is designed to speak text that has already been prepared by another model or application layer.
How much does Gemini 2.5 Flash TTS cost?
Google's listed standard pricing is $0.50 per 1 million input text tokens and $10.00 per 1 million output audio tokens. Batch pricing is $0.25 per 1 million input text tokens and $5.00 per 1 million output audio tokens. Actual costs depend on token usage and should be estimated with representative scripts.
What input and output modalities does Gemini 2.5 Flash TTS support?
Gemini 2.5 Flash TTS accepts text input and generates audio output. It supports single-speaker and multi-speaker synthesis, with prompts that can control characteristics such as tone, style, accent, and pacing. It does not accept audio, images, or video input and does not provide native text output.
What is Gemini 2.5 Flash TTS used for?
Gemini 2.5 Flash TTS is a specialized text-to-speech model for converting written text into spoken audio. It is suited to voice interfaces, read-aloud features, audiobook and podcast narration, simulated interviews, localized content, and batch audio production.
What is the model ID for Gemini 2.5 Flash TTS?
The canonical Gemini API model identifier is "gemini-2.5-flash-preview-tts". The model was released on May 20, 2025, and is currently classified as a preview model.


Sources 5
Provider

About Google DeepMind