Gemini 3.1 Flash Audio

Gemini 3.1 Flash TTS

by Google DeepMind · Legacy preview; currently accessible; no shutdown date announced

Gemini 3.1 Flash TTS is a legacy preview Google DeepMind model for generating expressive speech from text. It supports steerable delivery, audio tags, single- and multi-speaker synthesis, streaming, and batch processing through the Gemini API, Google AI Studio, and Vertex AI. Standard pricing is $1 per million text input tokens and $20 per million audio output tokens, while Google recommends newer Gemini 3.8 TTS models for new production use.

Speech Reasoning Coding
Gemini 3.1 Flash TTS converts written text into spoken audio with controls for delivery, pacing, tone, emphasis, and speaker interaction. It is designed for expressive narration, accessibility, scripted dialogue, and voice-interface prototypes. The model is available through the Gemini API, Google AI Studio, and Vertex AI, although Google now classifies it as a legacy preview model and recommends newer Gemini 3.8 TTS models for new production workloads.
Outputs

What Gemini 3.1 Flash TTS can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Batch API Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.1 Flash Audio
Model type Other
Context window 8K tokens
Maximum output 16K tokens
Knowledge cutoff January 2025
Release date 2026-02-26
Status Legacy preview; currently accessible; no shutdown date announced
Deprecation date 2026-02-26
Knowledge cutoff notes

Google DeepMind's model information page lists January 2025 as the knowledge cutoff. This cutoff applies to the underlying model and is not changed by external prompts, retrieval, or application-level context.

Model notes

Canonical API identifier is gemini-3.1-flash-tts-preview. Google classifies the model as a legacy preview model and recommends Gemini 3.8 Flash TTS or Gemini 3.8 Flash-Lite TTS for new and production workloads. The Gemini API model page lists 8,192 input tokens and 16,384 output tokens. Google DeepMind's model information page separately presents rounded 16k input and 32k output figures, so the API model-page limits are used here for implementation. The model supports single-speaker and multi-speaker synthesis, expressive audio tags, Google AI Studio, Gemini API, and Vertex AI. It does not support voice design or voice replication according to the TTS guide.

Cost

Model pricing

Input $1.00 per 1 million text tokens standard; $0.50 per 1 million text tokens batch
Output $20.00 per 1 million audio tokens standard; $10.00 per 1 million audio tokens batch
Model guide

Gemini 3.1 Flash TTS: Expressive Speech Generation in a Legacy Preview Model

Gemini 3.1 Flash TTS is a Google DeepMind preview text-to-speech model that generates expressive single-speaker or multi-speaker audio from text. It offers controllable pacing, tone, emphasis, emotion, and audio tags, but its legacy-preview status makes Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS more appropriate for new production deployments.

What is Gemini 3.1 Flash TTS?

Gemini 3.1 Flash TTS is a text-to-speech model from Google DeepMind. Its job is narrowly defined: it takes text as input and produces spoken audio. Unlike a general conversational audio model, it is not intended to understand speech, maintain a live voice conversation, or process images and video.

The model is identified in the Gemini API by the model ID gemini-3.1-flash-tts-preview. It is available through the Gemini API, Google AI Studio, and Vertex AI. Google lists it as a legacy preview model. That status matters when evaluating it for a new application: the model may remain accessible, but Google recommends Gemini 3.8 Flash TTS or Gemini 3.8 Flash-Lite TTS for new and production deployments.

Gemini 3.1 Flash TTS is best understood as a controllable speech renderer. The application supplies the words and can also describe how those words should be delivered. This makes it more suitable for expressive narration and scripted performance than for basic, unstyled read-aloud output.

Supported inputs and outputs

The documented TTS configuration accepts text input and returns audio output. The model does not provide image, video, or audio-understanding input in this configuration.

  • Input: Text
  • Output: Generated speech audio
  • Single speaker: Supported
  • Multiple speakers: Supported
  • Audio understanding: Not supported as a documented input capability
  • Image and video generation: Not supported

Audio output is the model's defining capability. It does not return ordinary text as its primary result, and it is not a speech-recognition model that transcribes recorded audio. Applications that need speech recognition, language-model reasoning, and speech synthesis must combine separate components rather than expecting Gemini 3.1 Flash TTS to perform all of those tasks.

How expressive speech control works

Gemini 3.1 Flash TTS supports steerable prompts, meaning the developer can describe the intended performance rather than merely submitting a block of text. Prompts can request changes in tone, pace, emphasis, emotion, and delivery. For example, a narration application may ask for a calm, measured explanation, while a fictional dialogue application may specify distinct delivery styles for different speakers.

The model also supports expressive audio tags. These tags can be used to guide performance details such as whispering, shouting, pauses, laughter, and other vocal events, subject to the model's prompting behavior. They are useful when the desired output depends on performance cues that would be difficult to express through punctuation alone.

Multi-speaker synthesis allows a scripted exchange to be generated as dialogue rather than as a single undifferentiated voice. This can help with prototypes for podcasts, educational conversations, interactive stories, and accessibility experiences. The supplied documentation does not indicate that the model supports voice design or voice replication, so it should not be treated as a system for cloning a particular person's voice.

Technical limits and API support

The Gemini API model page lists an 8,192-token input limit and a 16,384-token output limit. In this context, the output limit applies to generated audio-token output rather than to a conventional text response. Token counts therefore should not be interpreted as a direct character or word allowance for every language and speaking style.

SpecificationDocumented value
API model IDgemini-3.1-flash-tts-preview
Input limit8,192 tokens
Output limit16,384 audio tokens
Audio token accounting25 tokens per second of audio
Input modalityText
Output modalityAudio
Batch APISupported
StreamingSupported

The model page lists audio generation and batch API support. It does not support caching, function calling, search grounding, structured outputs, code execution, file search, Live API usage, or thinking. These limitations are important for architecture decisions. A voice application can use another model to generate or retrieve the script, then send the resulting text to Gemini 3.1 Flash TTS, but the TTS model itself is not a tool-using agent or a general-purpose reasoning model.

Gemini 3.1 Flash TTS pricing

Google's Gemini API pricing lists separate rates for text input and audio output. Standard pricing is $1.00 per 1 million text input tokens and $20.00 per 1 million audio output tokens. Batch pricing is lower: $0.50 per 1 million text input tokens and $10.00 per 1 million audio output tokens.

Processing modeText inputAudio output
Standard$1.00 per 1 million text tokens$20.00 per 1 million audio tokens
Batch$0.50 per 1 million text tokens$10.00 per 1 million audio tokens

Google also lists free-tier access. The effective audio cost depends primarily on the amount of generated speech, because audio output is priced separately from the text supplied to the model. Google states that audio is counted at 25 tokens per second. Batch processing can reduce cost when immediate responses are unnecessary, such as for pre-generating an audiobook chapter, accessibility library, or collection of scripted clips.

Strengths and trade-offs

The model's main strength is controllable expression. Many TTS systems can read text clearly, but Gemini 3.1 Flash TTS is aimed at cases where the manner of speaking is part of the output. Prompt-based control over pacing, tone, emphasis, emotion, and vocal events gives developers a way to adapt delivery without manually editing every recording.

  • Expressive delivery: Prompts and audio tags can guide tone, pacing, emphasis, pauses, and vocal events.
  • Dialogue generation: Single-speaker and multi-speaker synthesis support scripted conversations.
  • Speech-focused workflow: The model is purpose-built for turning prepared text into audio.
  • Integration options: Developers can use the Gemini API, Google AI Studio, or Vertex AI.
  • Low-latency positioning: The research describes the model as suitable for low-latency speech generation, while streaming is listed as supported.
  • Batch economics: Batch pricing is lower for workloads that can wait for processing.

Those strengths come with corresponding trade-offs. The model is not a complete voice assistant, because it does not provide speech recognition, function calling, search grounding, or live bidirectional audio interaction. It also lacks structured JSON output and code execution. A developer must supply surrounding application logic for text generation, conversation state, tool use, and audio input handling.

The legacy preview designation is another significant trade-off. Even if the model remains accessible and has useful expressive behavior, a new production system may be better served by a currently recommended model with a clearer support path. The supplied research does not provide benchmark results that would establish a numerical quality advantage over the newer models.

Reasoning, coding, and tool capabilities

Gemini 3.1 Flash TTS is not designed for independent reasoning or software development. Its editorial reasoning assessment is low because the model's role is speech synthesis, not because it is intended to answer reasoning questions. Similarly, its coding usefulness is limited to rendering text about code as speech; it is not a coding model.

The model does not support function calling, search grounding, code execution, file search, or structured outputs. It also does not support thinking. If an application needs to research information, call an external service, validate a response, or generate structured data, those steps must happen before or around the TTS request using other software or models.

Best use cases

Gemini 3.1 Flash TTS is a reasonable fit when the input script is known or generated elsewhere and the desired result is expressive speech. Suitable examples include:

  • Audiobook-style narration and short-form storytelling
  • Read-aloud features for accessibility applications
  • Scripted multi-speaker dialogue
  • Educational explanations with deliberately controlled pacing
  • Voice interfaces that use a separate language model for conversation and this model for speech rendering
  • Rapid speech-generation prototypes in Google AI Studio or Vertex AI
  • Batch creation of prerecorded audio clips

For a conversational assistant, the model can occupy the final speech-generation stage: another system can receive the user's request, reason about it, use tools, and produce a response script, while Gemini 3.1 Flash TTS turns that script into audio. This division of responsibilities is important because the TTS model itself does not perform the assistant's reasoning or tool calls.

When should you choose Gemini 3.1 Flash TTS?

Choose Gemini 3.1 Flash TTS when expressive text-to-speech is the immediate requirement, the available API behavior meets your needs, and you are building a prototype or a workload that can accept legacy-preview status. It is particularly relevant when delivery control, audio tags, multi-speaker output, and Google API access matter more than general-purpose model capabilities.

For a new production deployment, the supplied Google guidance points toward Gemini 3.8 Flash TTS or Gemini 3.8 Flash-Lite TTS instead. Gemini 3.8 Flash TTS is positioned for higher acoustic fidelity and expressive control, while Gemini 3.8 Flash-Lite TTS is positioned as the faster and more cost-efficient option for high-volume workloads. These recommendations create a practical choice: use the newer full model when speech fidelity and expression are priorities, or consider the Lite model when throughput and cost are more important.

A conventional speech-synthesis service may be more appropriate if the project requires a clearly supported production voice catalog, voice replication, or capabilities not documented for Gemini 3.1 Flash TTS. A general multimodal conversational model may be more suitable when the system must directly understand incoming audio, maintain a live two-way conversation, or combine speech with broader reasoning and tool use.

Status and migration considerations

Google lists Gemini 3.1 Flash TTS Preview as a legacy preview model. The supplied research gives its release date as February 26, 2026 and reports that no shutdown date has been announced. No shutdown date should not be interpreted as a long-term production guarantee.

Teams already testing the model should record the exact API identifier, pricing mode, audio settings, prompt conventions, and multi-speaker behavior used by their application. Before migrating, they should compare generated audio from the recommended Gemini 3.8 models using representative scripts, especially scripts containing speaker changes, emotional cues, pauses, or nonstandard pronunciation. Since the research does not provide benchmark scores, migration decisions should be validated with the application's own quality, latency, and cost measurements.

Overall, Gemini 3.1 Flash TTS is a specialized expressive speech generator rather than a general AI assistant. Its controllable narration and dialogue features make it useful for prototyping and selected speech workloads, but its preview legacy status, lack of tools and audio understanding, and the availability of newer recommended TTS models should be part of any deployment decision.


Answers to Frequently Asked Questions

Should developers use Gemini 3.1 Flash TTS for new production applications?
Gemini 3.1 Flash TTS is listed as a legacy preview model, so it may be suitable for prototypes or existing workloads that need expressive speech generation. For new production deployments, Google recommends Gemini 3.8 Flash TTS or Gemini 3.8 Flash-Lite TTS, depending on whether speech fidelity and expression or speed and cost efficiency are the priority.
What are the limitations of Gemini 3.1 Flash TTS?
Gemini 3.1 Flash TTS generates speech from text but does not provide speech recognition, audio understanding, image or video generation, function calling, search grounding, code execution, structured outputs, or live bidirectional conversation. Applications must use other models or software for reasoning, tool use, conversation management, and audio input processing.
How much does Gemini 3.1 Flash TTS cost?
Standard pricing is $1.00 per 1 million text input tokens and $20.00 per 1 million audio output tokens. Batch pricing is $0.50 per 1 million text input tokens and $10.00 per 1 million audio output tokens. Audio is counted at 25 tokens per second.
What is Gemini 3.1 Flash TTS?
Gemini 3.1 Flash TTS is a Google DeepMind text-to-speech model that converts text into generated speech audio. Its API model ID is gemini-3.1-flash-tts-preview, and it is available through the Gemini API, Google AI Studio, and Vertex AI.
What expressive speech features does Gemini 3.1 Flash TTS support?
The model supports prompt-based control over tone, pacing, emphasis, emotion, and delivery. It also supports expressive audio tags for cues such as whispering, shouting, pauses, and laughter, along with single-speaker and multi-speaker speech synthesis.


Sources 8
Provider

About Google DeepMind