What is Gemini 3.1 Flash TTS?
Gemini 3.1 Flash TTS is a text-to-speech model from Google DeepMind. Its job is narrowly defined: it takes text as input and produces spoken audio. Unlike a general conversational audio model, it is not intended to understand speech, maintain a live voice conversation, or process images and video.
The model is identified in the Gemini API by the model ID gemini-3.1-flash-tts-preview. It is available through the Gemini API, Google AI Studio, and Vertex AI. Google lists it as a legacy preview model. That status matters when evaluating it for a new application: the model may remain accessible, but Google recommends Gemini 3.8 Flash TTS or Gemini 3.8 Flash-Lite TTS for new and production deployments.
Gemini 3.1 Flash TTS is best understood as a controllable speech renderer. The application supplies the words and can also describe how those words should be delivered. This makes it more suitable for expressive narration and scripted performance than for basic, unstyled read-aloud output.
Supported inputs and outputs
The documented TTS configuration accepts text input and returns audio output. The model does not provide image, video, or audio-understanding input in this configuration.
- Input: Text
- Output: Generated speech audio
- Single speaker: Supported
- Multiple speakers: Supported
- Audio understanding: Not supported as a documented input capability
- Image and video generation: Not supported
Audio output is the model's defining capability. It does not return ordinary text as its primary result, and it is not a speech-recognition model that transcribes recorded audio. Applications that need speech recognition, language-model reasoning, and speech synthesis must combine separate components rather than expecting Gemini 3.1 Flash TTS to perform all of those tasks.
How expressive speech control works
Gemini 3.1 Flash TTS supports steerable prompts, meaning the developer can describe the intended performance rather than merely submitting a block of text. Prompts can request changes in tone, pace, emphasis, emotion, and delivery. For example, a narration application may ask for a calm, measured explanation, while a fictional dialogue application may specify distinct delivery styles for different speakers.
The model also supports expressive audio tags. These tags can be used to guide performance details such as whispering, shouting, pauses, laughter, and other vocal events, subject to the model's prompting behavior. They are useful when the desired output depends on performance cues that would be difficult to express through punctuation alone.
Multi-speaker synthesis allows a scripted exchange to be generated as dialogue rather than as a single undifferentiated voice. This can help with prototypes for podcasts, educational conversations, interactive stories, and accessibility experiences. The supplied documentation does not indicate that the model supports voice design or voice replication, so it should not be treated as a system for cloning a particular person's voice.
Technical limits and API support
The Gemini API model page lists an 8,192-token input limit and a 16,384-token output limit. In this context, the output limit applies to generated audio-token output rather than to a conventional text response. Token counts therefore should not be interpreted as a direct character or word allowance for every language and speaking style.
| Specification | Documented value |
|---|---|
| API model ID | gemini-3.1-flash-tts-preview |
| Input limit | 8,192 tokens |
| Output limit | 16,384 audio tokens |
| Audio token accounting | 25 tokens per second of audio |
| Input modality | Text |
| Output modality | Audio |
| Batch API | Supported |
| Streaming | Supported |
The model page lists audio generation and batch API support. It does not support caching, function calling, search grounding, structured outputs, code execution, file search, Live API usage, or thinking. These limitations are important for architecture decisions. A voice application can use another model to generate or retrieve the script, then send the resulting text to Gemini 3.1 Flash TTS, but the TTS model itself is not a tool-using agent or a general-purpose reasoning model.
Gemini 3.1 Flash TTS pricing
Google's Gemini API pricing lists separate rates for text input and audio output. Standard pricing is $1.00 per 1 million text input tokens and $20.00 per 1 million audio output tokens. Batch pricing is lower: $0.50 per 1 million text input tokens and $10.00 per 1 million audio output tokens.
| Processing mode | Text input | Audio output |
|---|---|---|
| Standard | $1.00 per 1 million text tokens | $20.00 per 1 million audio tokens |
| Batch | $0.50 per 1 million text tokens | $10.00 per 1 million audio tokens |
Google also lists free-tier access. The effective audio cost depends primarily on the amount of generated speech, because audio output is priced separately from the text supplied to the model. Google states that audio is counted at 25 tokens per second. Batch processing can reduce cost when immediate responses are unnecessary, such as for pre-generating an audiobook chapter, accessibility library, or collection of scripted clips.
Strengths and trade-offs
The model's main strength is controllable expression. Many TTS systems can read text clearly, but Gemini 3.1 Flash TTS is aimed at cases where the manner of speaking is part of the output. Prompt-based control over pacing, tone, emphasis, emotion, and vocal events gives developers a way to adapt delivery without manually editing every recording.
- Expressive delivery: Prompts and audio tags can guide tone, pacing, emphasis, pauses, and vocal events.
- Dialogue generation: Single-speaker and multi-speaker synthesis support scripted conversations.
- Speech-focused workflow: The model is purpose-built for turning prepared text into audio.
- Integration options: Developers can use the Gemini API, Google AI Studio, or Vertex AI.
- Low-latency positioning: The research describes the model as suitable for low-latency speech generation, while streaming is listed as supported.
- Batch economics: Batch pricing is lower for workloads that can wait for processing.
Those strengths come with corresponding trade-offs. The model is not a complete voice assistant, because it does not provide speech recognition, function calling, search grounding, or live bidirectional audio interaction. It also lacks structured JSON output and code execution. A developer must supply surrounding application logic for text generation, conversation state, tool use, and audio input handling.
The legacy preview designation is another significant trade-off. Even if the model remains accessible and has useful expressive behavior, a new production system may be better served by a currently recommended model with a clearer support path. The supplied research does not provide benchmark results that would establish a numerical quality advantage over the newer models.
Reasoning, coding, and tool capabilities
Gemini 3.1 Flash TTS is not designed for independent reasoning or software development. Its editorial reasoning assessment is low because the model's role is speech synthesis, not because it is intended to answer reasoning questions. Similarly, its coding usefulness is limited to rendering text about code as speech; it is not a coding model.
The model does not support function calling, search grounding, code execution, file search, or structured outputs. It also does not support thinking. If an application needs to research information, call an external service, validate a response, or generate structured data, those steps must happen before or around the TTS request using other software or models.
Best use cases
Gemini 3.1 Flash TTS is a reasonable fit when the input script is known or generated elsewhere and the desired result is expressive speech. Suitable examples include:
- Audiobook-style narration and short-form storytelling
- Read-aloud features for accessibility applications
- Scripted multi-speaker dialogue
- Educational explanations with deliberately controlled pacing
- Voice interfaces that use a separate language model for conversation and this model for speech rendering
- Rapid speech-generation prototypes in Google AI Studio or Vertex AI
- Batch creation of prerecorded audio clips
For a conversational assistant, the model can occupy the final speech-generation stage: another system can receive the user's request, reason about it, use tools, and produce a response script, while Gemini 3.1 Flash TTS turns that script into audio. This division of responsibilities is important because the TTS model itself does not perform the assistant's reasoning or tool calls.
When should you choose Gemini 3.1 Flash TTS?
Choose Gemini 3.1 Flash TTS when expressive text-to-speech is the immediate requirement, the available API behavior meets your needs, and you are building a prototype or a workload that can accept legacy-preview status. It is particularly relevant when delivery control, audio tags, multi-speaker output, and Google API access matter more than general-purpose model capabilities.
For a new production deployment, the supplied Google guidance points toward Gemini 3.8 Flash TTS or Gemini 3.8 Flash-Lite TTS instead. Gemini 3.8 Flash TTS is positioned for higher acoustic fidelity and expressive control, while Gemini 3.8 Flash-Lite TTS is positioned as the faster and more cost-efficient option for high-volume workloads. These recommendations create a practical choice: use the newer full model when speech fidelity and expression are priorities, or consider the Lite model when throughput and cost are more important.
A conventional speech-synthesis service may be more appropriate if the project requires a clearly supported production voice catalog, voice replication, or capabilities not documented for Gemini 3.1 Flash TTS. A general multimodal conversational model may be more suitable when the system must directly understand incoming audio, maintain a live two-way conversation, or combine speech with broader reasoning and tool use.
Status and migration considerations
Google lists Gemini 3.1 Flash TTS Preview as a legacy preview model. The supplied research gives its release date as February 26, 2026 and reports that no shutdown date has been announced. No shutdown date should not be interpreted as a long-term production guarantee.
Teams already testing the model should record the exact API identifier, pricing mode, audio settings, prompt conventions, and multi-speaker behavior used by their application. Before migrating, they should compare generated audio from the recommended Gemini 3.8 models using representative scripts, especially scripts containing speaker changes, emotional cues, pauses, or nonstandard pronunciation. Since the research does not provide benchmark scores, migration decisions should be validated with the application's own quality, latency, and cost measurements.
Overall, Gemini 3.1 Flash TTS is a specialized expressive speech generator rather than a general AI assistant. Its controllable narration and dialogue features make it useful for prototyping and selected speech workloads, but its preview legacy status, lack of tools and audio understanding, and the availability of newer recommended TTS models should be part of any deployment decision.

