Gemini 3.8

Gemini 3.8 Flash TTS

by Google DeepMind · Generally available (GA); no shutdown date announced

Google DeepMind's generally available Gemini 3.8 Flash TTS converts text into expressive audio for narration, audiobooks, voice acting, multi-speaker dialogue, regional accents, voice design, and voice replication. It supports more than 130 languages, streaming, caching, and batch processing, with an 8,192-token context limit and usage-based audio pricing.

Speech Reasoning Coding
Gemini 3.8 Flash TTS is Google DeepMind's dedicated text-to-speech model for turning written text into expressive spoken audio. It is aimed at applications such as narration, audiobooks, voice acting, dialogue, regional pronunciation, and other production workflows where voice quality and consistency matter more than general-purpose reasoning.
Outputs

What Gemini 3.8 Flash TTS can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Prompt caching Batch API Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.8
Model type Other
Context window 8K tokens
Maximum output 16K tokens
Release date September 22, 2026
Status Generally available (GA); no shutdown date announced
Knowledge cutoff notes

Google's current model documentation does not publish a separate knowledge cutoff for this text-to-speech model. The model is intended for text-to-audio synthesis rather than knowledge-intensive conversational answering.

Model notes

Canonical model ID: gemini-3.8-flash-tts. The model accepts text-only input and produces audio-only output. It supports single-speaker and multi-speaker synthesis, more than 130 languages, prebuilt voices, an extended voice library, custom voice design, and voice replication. Unary requests return WAV audio by default at 24 kHz mono 16-bit PCM; streaming requests return headerless audio/l16 PCM chunks by default. Supported alternative output encodings include audio/l16, audio/mulaw, and audio/alaw. Sustained delivery instructions belong in speech_metadata, while point-in-time vocal events can use inline tags such as <laugh>, <sigh>, and <short pause>. Batch API, Flex inference, Priority inference, and context caching are supported. Editorial scores are comparative estimates for this specialist TTS model and are not vendor benchmarks.

Cost

Model pricing

Input $0.50 per 1 million text input tokens through December 31, 2026; $1.00 per 1 million text input tokens from January 1, 2027
Output $9.00 per 1 million audio output tokens through December 31, 2026; $18.00 per 1 million audio output tokens from January 1, 2027. Equivalent to $0.00225 per 10 seconds of audio through December 31, 2026 and $0.0045 per 10 seconds thereafter.
Model guide

Gemini 3.8 Flash TTS: Expressive Text-to-Speech for Long-Form Audio

Google DeepMind's Gemini 3.8 Flash TTS is a generally available text-to-speech model designed for high-fidelity, expressive audio generation. It converts text into speech-only output and supports multi-speaker dialogue, long-form narration, regional accents, voice design, voice replication, streaming, caching, and batch processing.

What is Gemini 3.8 Flash TTS?

Gemini 3.8 Flash TTS is a generally available text-to-speech model from Google DeepMind. Its input is text and its output is audio, so it is not a conversational language model intended to answer questions, write code, search the web, or return structured data. Instead, it specializes in synthesizing speech from written instructions and scripts.

The model is positioned as Google's flagship creative TTS model in the supplied product research. Its focus is acoustic fidelity, expressive delivery, dialect coverage, and stable voice identity over long-form or multi-turn generation. In practical terms, that makes it more relevant to narration and voice production than to ordinary text-generation tasks.

The canonical model identifier is gemini-3.8-flash-tts. The model was released on September 22, 2026, and is listed as generally available, with no announced shutdown date.

Main capabilities

Gemini 3.8 Flash TTS can generate both single-speaker and multi-speaker speech. Multi-speaker synthesis is useful for scripted conversations, interviews, audio drama, training simulations, and dialogue-heavy educational content. The model can also work with more than 130 languages, making it suitable for multilingual narration and localized voice experiences.

Its supported voice features include prebuilt voices, an extended voice library, custom voice design, and voice replication. Voice design allows a user to specify the desired characteristics of a voice, while voice replication is intended to reproduce the identity or style of a reference voice where the relevant workflow and permissions allow it. The supplied research does not specify the exact reference-audio requirements or consent controls, so those details should be checked in the current implementation documentation before deployment.

Speech direction can include sustained delivery instructions through speech_metadata. Point-in-time vocal events can be represented with inline tags such as <laugh>, <sigh>, and <short pause>. This distinction is useful: ongoing characteristics such as tone or delivery belong in broader speech instructions, while short events can be placed at specific points in a script.

Audio output, formats, and streaming

Unary requests return WAV audio by default at 24 kHz, mono, 16-bit PCM. Streaming requests return headerless audio/L16 PCM chunks by default. The documented alternative encodings are audio/L16, audio/mulaw, and audio/alaw.

Streaming is important when an application needs to begin playing speech before the complete response has been generated. It can reduce perceived waiting time for interactive narration or dialogue, although the supplied research does not provide a measured latency benchmark. For file-based production, a complete WAV response may be more convenient because it includes the expected audio container and can be saved directly for later editing or distribution.

The model supports caching and batch processing. Context caching can be useful when the same instructions or recurring material are reused across multiple generations. Batch processing is better suited to non-urgent workloads such as producing many chapters, localized versions, or large sets of voice assets. The model also supports Flex inference and Priority inference, although the supplied information does not define their exact service-level or pricing differences.

Technical specifications

SpecificationGemini 3.8 Flash TTS
ProviderGoogle DeepMind
AvailabilityGenerally available
Model IDgemini-3.8-flash-tts
InputText only
OutputAudio only
Context length8,192 tokens
Maximum output tokens16,384
StreamingSupported
CachingSupported
Batch APISupported
Fine-tuningNot supported

The 8,192-token context length limits how much text and instruction material can be supplied in a request. For long books, scripts, or serialized content, the practical workflow may require splitting the source into chapters or sections. Splitting also makes it easier to regenerate a problematic passage without recreating an entire production. The maximum output is listed as 16,384 tokens, but audio duration will depend on the amount of spoken content and the model's tokenization.

Pricing

According to the supplied Google AI pricing research, text input costs $0.50 per 1 million text input tokens through December 31, 2026. From January 1, 2027, the listed input price is $1.00 per 1 million text input tokens.

Audio output costs $9.00 per 1 million audio output tokens through December 31, 2026, increasing to $18.00 per 1 million audio output tokens from January 1, 2027. The supplied pricing notes give an equivalent output cost of approximately $0.00225 per 10 seconds of audio before January 1, 2027, and $0.0045 per 10 seconds thereafter.

These prices are usage-based rather than a recurring consumer subscription price. Actual project costs depend primarily on generated audio volume, repeated requests, retries, and the use of batch, cached, Flex, or Priority processing. Developers should confirm the current provider pricing page before committing to a production budget because the supplied figures include a dated promotional or transitional pricing period.

Strengths and trade-offs

The model's main strength is specialization. A dedicated TTS model can be a better fit than a general-purpose text model when the important outcome is natural, expressive, consistent speech. The documented support for multi-speaker output, regional accents, difficult pronunciations, long-form narration, voice design, and voice replication gives it a broad production-oriented feature set.

Its audio-only design is also a limitation. Gemini 3.8 Flash TTS does not accept image, audio, or video input according to the supplied specification. It is not a speech-recognition model, so it should not be selected for transcribing recordings. It also does not provide text output, image or video generation, web search, tool calling, structured JSON generation, or general reasoning.

There is no fine-tuning support listed. Users who need model-level training on a proprietary voice or a highly specialized pronunciation set cannot rely on fine-tuning as a documented option. Instead, they would need to use the available voice controls, prompting, voice design, replication features, and application-side preparation described by the provider.

The editorial research rates the model highly for speed and cost relative to its specialist TTS role, with comparative scores of 8 for speed and 7 for cost. These are editorial estimates, not provider-published benchmarks. The same research gives reasoning and coding scores of 1, but those scores should be understood as suitability ratings: reasoning and coding are outside the model's intended purpose, rather than evidence that it is a general-purpose model with weak reasoning.

When to choose Gemini 3.8 Flash TTS

Choose Gemini 3.8 Flash TTS when the central requirement is generated speech rather than text intelligence. It is a strong candidate for:

  • Audiobooks and long-form narration that need a stable voice across many sections.
  • Studio-quality voice acting for scripted dialogue or audio drama.
  • Multi-speaker conversations, interviews, simulations, and training material.
  • Localized narration requiring support for more than 130 languages or regional accents.
  • Voice experiences that need designed voices, replicated voices, expressive delivery, or controlled vocal events.
  • High-volume production pipelines that can benefit from caching and batch processing.
  • Interactive applications where streaming audio can reduce the time before playback begins.

It may be especially appropriate when consistency and expressiveness matter more than having one model perform several unrelated tasks. For example, an application can use a separate system to write or organize a script and then send the approved text to Gemini 3.8 Flash TTS for rendering.

When another option may be more appropriate

Use a general-purpose language model instead when the primary task is writing, summarization, question answering, coding, reasoning, or structured data generation. Gemini 3.8 Flash TTS is not designed to perform those jobs and does not return text output or structured JSON.

A speech-to-speech or audio-capable model may be more suitable when the application must listen to a user's voice, analyze an audio recording, or maintain an audio-to-audio interaction. The supplied research specifically lists Gemini 3.8 Flash TTS as unsuitable for speech recognition and real-time Live API audio-to-audio interaction.

A different TTS service may be preferable if the project requires fine-tuning, a particular deployment environment, a provider-specific voice catalog, or contract terms not documented for this model. Likewise, teams with very large scripts should account for the 8,192-token context limit and design a reliable chunking and continuity strategy before production.

Bottom line

Gemini 3.8 Flash TTS is a focused audio-generation model rather than an all-purpose Gemini assistant. Its value comes from expressive speech synthesis, multi-speaker dialogue, multilingual coverage, voice customization, stable long-form delivery, streaming, and production-oriented processing options. The most important constraints are its text-only input, audio-only output, 8,192-token context limit, lack of fine-tuning, and lack of general reasoning, coding, tool-use, or search capabilities.

For developers and creators who already have text scripts and need polished spoken audio, it offers a clear specialist workflow. For applications that must understand audio, generate text, call tools, or reason over external information, it should be treated as one component in a larger system rather than the primary model.


Answers to Frequently Asked Questions

How much does Gemini 3.8 Flash TTS cost?
According to the supplied pricing research, text input costs $0.50 per 1 million tokens and audio output costs $9.00 per 1 million audio output tokens through December 31, 2026. From January 1, 2027, the listed prices increase to $1.00 per 1 million text input tokens and $18.00 per 1 million audio output tokens. Developers should confirm current pricing before production use.
What are the limitations of Gemini 3.8 Flash TTS?
Gemini 3.8 Flash TTS accepts text only and produces audio only. It is not designed for speech recognition, transcription, question answering, coding, reasoning, web search, tool calling, structured JSON generation, or audio-to-audio interaction. It has an 8,192-token context limit and does not support fine-tuning.
What audio formats and output settings does Gemini 3.8 Flash TTS support?
Unary requests return WAV audio by default at 24 kHz, mono, 16-bit PCM. Streaming requests return headerless audio/L16 PCM chunks by default. The documented alternative encodings are audio/L16, audio/mulaw, and audio/alaw.
What is Gemini 3.8 Flash TTS used for?
Gemini 3.8 Flash TTS is a specialized text-to-speech model for generating expressive audio from written scripts. It is suited to audiobooks, long-form narration, scripted dialogue, audio drama, interviews, training simulations, multilingual voice experiences, and interactive applications that need streamed speech.
What are the main capabilities of Gemini 3.8 Flash TTS?
The model supports single-speaker and multi-speaker speech, more than 130 languages, regional accents, prebuilt voices, voice design, voice replication, expressive delivery instructions, and inline vocal events such as , , and . It also supports streaming, caching, and batch processing.


Sources 5
Provider

About Google DeepMind