What is Gemini 3.8 Flash TTS?
Gemini 3.8 Flash TTS is a generally available text-to-speech model from Google DeepMind. Its input is text and its output is audio, so it is not a conversational language model intended to answer questions, write code, search the web, or return structured data. Instead, it specializes in synthesizing speech from written instructions and scripts.
The model is positioned as Google's flagship creative TTS model in the supplied product research. Its focus is acoustic fidelity, expressive delivery, dialect coverage, and stable voice identity over long-form or multi-turn generation. In practical terms, that makes it more relevant to narration and voice production than to ordinary text-generation tasks.
The canonical model identifier is gemini-3.8-flash-tts. The model was released on September 22, 2026, and is listed as generally available, with no announced shutdown date.
Main capabilities
Gemini 3.8 Flash TTS can generate both single-speaker and multi-speaker speech. Multi-speaker synthesis is useful for scripted conversations, interviews, audio drama, training simulations, and dialogue-heavy educational content. The model can also work with more than 130 languages, making it suitable for multilingual narration and localized voice experiences.
Its supported voice features include prebuilt voices, an extended voice library, custom voice design, and voice replication. Voice design allows a user to specify the desired characteristics of a voice, while voice replication is intended to reproduce the identity or style of a reference voice where the relevant workflow and permissions allow it. The supplied research does not specify the exact reference-audio requirements or consent controls, so those details should be checked in the current implementation documentation before deployment.
Speech direction can include sustained delivery instructions through speech_metadata. Point-in-time vocal events can be represented with inline tags such as <laugh>, <sigh>, and <short pause>. This distinction is useful: ongoing characteristics such as tone or delivery belong in broader speech instructions, while short events can be placed at specific points in a script.
Audio output, formats, and streaming
Unary requests return WAV audio by default at 24 kHz, mono, 16-bit PCM. Streaming requests return headerless audio/L16 PCM chunks by default. The documented alternative encodings are audio/L16, audio/mulaw, and audio/alaw.
Streaming is important when an application needs to begin playing speech before the complete response has been generated. It can reduce perceived waiting time for interactive narration or dialogue, although the supplied research does not provide a measured latency benchmark. For file-based production, a complete WAV response may be more convenient because it includes the expected audio container and can be saved directly for later editing or distribution.
The model supports caching and batch processing. Context caching can be useful when the same instructions or recurring material are reused across multiple generations. Batch processing is better suited to non-urgent workloads such as producing many chapters, localized versions, or large sets of voice assets. The model also supports Flex inference and Priority inference, although the supplied information does not define their exact service-level or pricing differences.
Technical specifications
| Specification | Gemini 3.8 Flash TTS |
|---|---|
| Provider | Google DeepMind |
| Availability | Generally available |
| Model ID | gemini-3.8-flash-tts |
| Input | Text only |
| Output | Audio only |
| Context length | 8,192 tokens |
| Maximum output tokens | 16,384 |
| Streaming | Supported |
| Caching | Supported |
| Batch API | Supported |
| Fine-tuning | Not supported |
The 8,192-token context length limits how much text and instruction material can be supplied in a request. For long books, scripts, or serialized content, the practical workflow may require splitting the source into chapters or sections. Splitting also makes it easier to regenerate a problematic passage without recreating an entire production. The maximum output is listed as 16,384 tokens, but audio duration will depend on the amount of spoken content and the model's tokenization.
Pricing
According to the supplied Google AI pricing research, text input costs $0.50 per 1 million text input tokens through December 31, 2026. From January 1, 2027, the listed input price is $1.00 per 1 million text input tokens.
Audio output costs $9.00 per 1 million audio output tokens through December 31, 2026, increasing to $18.00 per 1 million audio output tokens from January 1, 2027. The supplied pricing notes give an equivalent output cost of approximately $0.00225 per 10 seconds of audio before January 1, 2027, and $0.0045 per 10 seconds thereafter.
These prices are usage-based rather than a recurring consumer subscription price. Actual project costs depend primarily on generated audio volume, repeated requests, retries, and the use of batch, cached, Flex, or Priority processing. Developers should confirm the current provider pricing page before committing to a production budget because the supplied figures include a dated promotional or transitional pricing period.
Strengths and trade-offs
The model's main strength is specialization. A dedicated TTS model can be a better fit than a general-purpose text model when the important outcome is natural, expressive, consistent speech. The documented support for multi-speaker output, regional accents, difficult pronunciations, long-form narration, voice design, and voice replication gives it a broad production-oriented feature set.
Its audio-only design is also a limitation. Gemini 3.8 Flash TTS does not accept image, audio, or video input according to the supplied specification. It is not a speech-recognition model, so it should not be selected for transcribing recordings. It also does not provide text output, image or video generation, web search, tool calling, structured JSON generation, or general reasoning.
There is no fine-tuning support listed. Users who need model-level training on a proprietary voice or a highly specialized pronunciation set cannot rely on fine-tuning as a documented option. Instead, they would need to use the available voice controls, prompting, voice design, replication features, and application-side preparation described by the provider.
The editorial research rates the model highly for speed and cost relative to its specialist TTS role, with comparative scores of 8 for speed and 7 for cost. These are editorial estimates, not provider-published benchmarks. The same research gives reasoning and coding scores of 1, but those scores should be understood as suitability ratings: reasoning and coding are outside the model's intended purpose, rather than evidence that it is a general-purpose model with weak reasoning.
When to choose Gemini 3.8 Flash TTS
Choose Gemini 3.8 Flash TTS when the central requirement is generated speech rather than text intelligence. It is a strong candidate for:
- Audiobooks and long-form narration that need a stable voice across many sections.
- Studio-quality voice acting for scripted dialogue or audio drama.
- Multi-speaker conversations, interviews, simulations, and training material.
- Localized narration requiring support for more than 130 languages or regional accents.
- Voice experiences that need designed voices, replicated voices, expressive delivery, or controlled vocal events.
- High-volume production pipelines that can benefit from caching and batch processing.
- Interactive applications where streaming audio can reduce the time before playback begins.
It may be especially appropriate when consistency and expressiveness matter more than having one model perform several unrelated tasks. For example, an application can use a separate system to write or organize a script and then send the approved text to Gemini 3.8 Flash TTS for rendering.
When another option may be more appropriate
Use a general-purpose language model instead when the primary task is writing, summarization, question answering, coding, reasoning, or structured data generation. Gemini 3.8 Flash TTS is not designed to perform those jobs and does not return text output or structured JSON.
A speech-to-speech or audio-capable model may be more suitable when the application must listen to a user's voice, analyze an audio recording, or maintain an audio-to-audio interaction. The supplied research specifically lists Gemini 3.8 Flash TTS as unsuitable for speech recognition and real-time Live API audio-to-audio interaction.
A different TTS service may be preferable if the project requires fine-tuning, a particular deployment environment, a provider-specific voice catalog, or contract terms not documented for this model. Likewise, teams with very large scripts should account for the 8,192-token context limit and design a reliable chunking and continuity strategy before production.
Bottom line
Gemini 3.8 Flash TTS is a focused audio-generation model rather than an all-purpose Gemini assistant. Its value comes from expressive speech synthesis, multi-speaker dialogue, multilingual coverage, voice customization, stable long-form delivery, streaming, and production-oriented processing options. The most important constraints are its text-only input, audio-only output, 8,192-token context limit, lack of fine-tuning, and lack of general reasoning, coding, tool-use, or search capabilities.
For developers and creators who already have text scripts and need polished spoken audio, it offers a clear specialist workflow. For applications that must understand audio, generate text, call tools, or reason over external information, it should be treated as one component in a larger system rather than the primary model.

