Gemini 3.8

Gemini 3.8 Flash-Lite TTS

by Google DeepMind · Generally available

Google's Gemini 3.8 Flash-Lite TTS is a generally available text-to-speech model optimized for high throughput, low latency, and cost efficiency. It converts text into audio, supports more than 100 languages, single- and multi-speaker synthesis, voice design, voice replication, and inline vocal events. It is best suited to scalable speech production rather than reasoning, transcription, tool use, or premium studio narration.

Speech Reasoning Coding
Gemini 3.8 Flash-Lite TTS is designed for applications that need to generate speech repeatedly and quickly without the cost or latency associated with a premium narration model. It accepts text and returns audio, making it suitable for read-aloud features, large-scale content production, voice-agent responses, everyday assistant speech, and voice replication workflows.
Outputs

What Gemini 3.8 Flash-Lite TTS can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Prompt caching Batch API Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.8
Model type Other
Context window 8K tokens
Maximum output 16K tokens
Release date 2026-09-22
Status Generally available
Model notes

Canonical model ID is gemini-3.8-flash-lite-tts. Generally available since September 22, 2026 and recommended as the successor to gemini-3.1-flash-tts-preview. Supports single-speaker and multi-speaker TTS, Voice design, Voice replication, prebuilt voices, and the Extended Voice Library. Input text is treated as a verbatim transcript; sustained style and speaker instructions belong in speech_metadata. Unary requests return WAV audio by default. Prices are introductory through December 31, 2026 and are scheduled to increase on January 1, 2027. Audio billing uses 25 tokens per second of audio.

Cost

Model pricing

Input $0.50 per 1 million text tokens through December 31, 2026; $1.00 per 1 million text tokens starting January 1, 2027. Batch and Flex input: $0.25 through December 31, 2026; $0.50 starting January 1, 2027.
Output $6.00 per 1 million audio tokens through December 31, 2026; $12.00 per 1 million audio tokens starting January 1, 2027. Standard equivalent: approximately $0.0015 per 10 seconds of audio through December 31, 2026.
Model guide

Gemini 3.8 Flash-Lite TTS: Cost-Efficient Speech for High-Volume Audio

Gemini 3.8 Flash-Lite TTS is Google's generally available text-to-speech model for fast, scalable, and cost-conscious audio generation. It converts text into spoken audio, supports single- and multi-speaker synthesis, voice design, voice replication, structured speech metadata, inline vocal events, and more than 100 languages. Its main trade-off is that it prioritizes throughput and price over the highest level of expressive fidelity.

What is Gemini 3.8 Flash-Lite TTS?

Gemini 3.8 Flash-Lite TTS is a text-to-speech model from Google. Its canonical Gemini API model identifier is gemini-3.8-flash-lite-tts. Unlike a general-purpose Gemini language model, it is focused on turning supplied text into spoken audio.

Google positions the model as the workhorse option in the Gemini 3.8 TTS family. The Flash-Lite designation indicates a focus on high throughput, low latency, and cost efficiency. It uses the same API schema and structured prompting approach as Gemini 3.8 Flash TTS, but is intended for workloads where economical, scalable speech generation matters more than maximum acoustic fidelity or highly nuanced studio-style performance.

The model was released as generally available on September 22, 2026. The information and introductory pricing described here reflect the supplied Google documentation and pricing details for September 2026.

How the model handles input and output

Flash-Lite TTS accepts text input and produces audio output. The supplied text is treated as the transcript to be spoken. Instructions about sustained characteristics such as tone, pace, whispering, or delivery style should be provided through structured speech_metadata rather than inserted into the transcript as ordinary stage directions.

The model supports both single-speaker and multi-speaker generation. For a multi-speaker request, speaker labels should be specified in speech metadata for every turn. This makes it possible to generate dialogue, narrated conversations, or other content that requires more than one voice.

  • Input: Text
  • Output: Spoken audio
  • Speakers: Single-speaker and multi-speaker synthesis
  • Languages: More than 100
  • Default unary format: WAV
  • Other audio formats: Raw PCM, mu-law, and A-law are available as options
  • Input limit: 8,192 tokens
  • Output serving limit: 16,384 tokens

Unary requests return WAV audio by default. Applications migrating from older Gemini TTS preview models should take care not to add a second WAV header to the returned bytes.

Voice and speech capabilities

The model includes several capabilities beyond basic text-to-speech. It can use prebuilt voices, the Extended Voice Library, custom Voice design personas, and Voice replication. Voice design allows an application to describe the kind of voice it wants, while Voice replication is intended for workflows that need a voice reproduced from an appropriate source or configuration.

Structured speech metadata can describe speakers and delivery style. The model also supports inline vocal events, including laughter, sighs, coughs, breaths, and short pauses. These events can be represented with tags such as <laugh>, <sigh>, and <short pause>. They are useful when a script needs some conversational texture rather than a completely uniform read.

These controls do not turn Flash-Lite TTS into a general audio-production or music-generation system. Its purpose remains speech synthesis from text.

Where Gemini 3.8 Flash-Lite TTS fits best

Flash-Lite TTS is most useful when an application needs speech generated frequently, with predictable API-based processing and controlled cost. Suitable examples include:

  • High-volume production of narrated or read-aloud content
  • Accessibility features that read articles, documents, or interface text aloud
  • Low-latency responses in conversational voice-agent cascades
  • Everyday assistant responses where speech must be generated quickly
  • Multi-speaker dialogue and scripted conversational content
  • Voice replication workflows that require scalable synthesis
  • Applications that need speech in a broad range of languages

It is particularly attractive for pipelines in which a separate system handles conversation logic, retrieval, or reasoning and Flash-Lite TTS produces the final spoken response. The model itself should not be selected as the reasoning component of that pipeline: it is a speech-generation model, not a general-purpose chat or reasoning model.

Pricing, limits, and availability

Google lists the following standard paid pricing for Gemini 3.8 Flash-Lite TTS through December 31, 2026:

Usage typePrice
Text input$0.50 per 1 million text tokens
Audio output$6.00 per 1 million audio tokens
Standard output equivalentApproximately $0.0015 per 10 seconds of audio

Beginning January 1, 2027, the listed standard prices are scheduled to increase to $1.00 per 1 million text input tokens and $12.00 per 1 million audio output tokens. Batch and Flex input is listed at $0.25 per 1 million tokens through December 31, 2026, increasing to $0.50 from January 1, 2027.

Audio billing uses 25 tokens per second of audio. The exact amount paid for a request therefore depends on the generated duration as well as the amount of text submitted. The introductory prices are date-limited, so production cost estimates should account for the scheduled 2027 increase rather than assuming the temporary rates will continue.

The model has an 8,192-token input limit and a 16,384-token output serving limit. Google lists caching, the Batch API, Flex inference, and Priority inference as supported. It is generally available rather than a preview-only model.

Supported and unsupported API features

Flash-Lite TTS supports speech generation features, but it does not expose the broader tool and output capabilities associated with some general-purpose Gemini models. The supplied specifications indicate that the following are not supported for this exact model:

  • Function calling
  • Code execution
  • File search
  • Live API
  • Search grounding
  • Google Maps grounding
  • Structured outputs
  • Thinking
  • URL context
  • Image generation

It also does not accept audio, image, or video input according to the supplied model specifications. Its multimodal capability is therefore limited to producing audio from text, rather than accepting arbitrary media for analysis or participating in audio-to-audio conversations.

Speed, cost, and capability trade-offs

The central choice behind Flash-Lite TTS is a production trade-off. Its design favors speed and cost control over the most detailed expressive performance. For a service generating thousands or millions of spoken responses, lower per-request cost and low latency can matter more than subtle improvements in acting, acoustic detail, or dialect-specific performance.

A premium speech option may be more appropriate when the primary requirement is highly expressive narration, maximum acoustic fidelity, extensive dialect coverage, or studio-quality voice performance. The supplied research specifically positions Gemini 3.8 Flash TTS as the more suitable sibling when those qualities are more important. Flash-Lite TTS is the better fit when the workload is repetitive, scalable, and price-sensitive.

The model is also a poor fit for tasks that require reasoning, coding, transcription, image generation, video generation, or structured JSON responses. Those functions should be handled by a different model or service before the resulting text is passed to Flash-Lite TTS.

Implementation guidance

Applications should separate transcript content from speech instructions. For example, the words intended to be spoken should remain in the text transcript, while persistent directions about pace, tone, or delivery should be represented through speech metadata. In multi-speaker requests, every turn should identify its speaker through the same structured mechanism.

Inline event tags are appropriate for momentary vocal actions. A script can include tags for a laugh, sigh, breath, cough, or short pause where supported. These should be used for local events rather than as a replacement for sustained style instructions.

Because unary responses are WAV by default, audio-handling code should preserve the returned format correctly. Adding another WAV header can corrupt the output. If an application needs raw PCM, mu-law, or A-law, it should request and process the selected format consistently throughout the pipeline.

When should you choose Gemini 3.8 Flash-Lite TTS?

Choose Gemini 3.8 Flash-Lite TTS when you need text-to-speech that is fast, comparatively inexpensive, and suitable for repeated production use. It is a strong candidate for read-aloud applications, high-volume narration, voice-agent cascades, everyday assistant speech, and voice workflows that benefit from single- or multi-speaker generation.

Choose another option when the project depends on premium studio narration, the broadest language or dialect coverage, maximum expressive control, audio-to-audio interaction, transcription, reasoning, or tool use. In particular, the flagship Gemini 3.8 Flash TTS model may be more appropriate when expressive fidelity is more important than the Flash-Lite model's throughput and cost advantages.

Overall assessment

Gemini 3.8 Flash-Lite TTS is a focused speech model rather than a general AI assistant. Its value comes from combining text-to-speech generation, multi-speaker support, voice design, voice replication, vocal events, broad language coverage, and production-oriented pricing in a model designed for scale.

Its limitations are equally important: it does not reason, call tools, analyze uploaded media, return structured JSON, or provide Live API audio-to-audio interaction. For organizations that need a dependable speech layer after another system has produced the text, those restrictions are usually a clear boundary rather than a disadvantage. The model's practical appeal is greatest when fast and economical speech generation is the requirement.


Answers to Frequently Asked Questions

What are the main limitations of Gemini 3.8 Flash-Lite TTS?
Gemini 3.8 Flash-Lite TTS is a speech-generation model rather than a general-purpose AI assistant. It does not support reasoning, function calling, code execution, file search, Live API, grounding, structured outputs, image generation, or audio, image, and video inputs. It is also optimized for throughput and cost rather than premium studio-level expressiveness.
What audio formats and speaker options does Gemini 3.8 Flash-Lite TTS support?
The model supports single-speaker and multi-speaker synthesis in more than 100 languages. Unary requests return WAV audio by default, with raw PCM, mu-law, and A-law available as options. It also supports prebuilt voices, voice design, voice replication, structured speech metadata, and vocal events such as laughter, sighs, breaths, coughs, and short pauses.
How much does Gemini 3.8 Flash-Lite TTS cost?
Through December 31, 2026, standard paid pricing is $0.50 per 1 million text input tokens and $6.00 per 1 million audio output tokens, equivalent to approximately $0.0015 per 10 seconds of audio. From January 1, 2027, the listed prices are scheduled to increase to $1.00 per 1 million text tokens and $12.00 per 1 million audio tokens.
What is Gemini 3.8 Flash-Lite TTS used for?
Gemini 3.8 Flash-Lite TTS is designed for fast, cost-efficient text-to-speech generation at scale. Common uses include high-volume narration, accessibility readers, conversational voice agents, everyday assistant responses, multi-speaker dialogue, and voice replication workflows.
What is the model identifier for Gemini 3.8 Flash-Lite TTS?
The canonical Gemini API model identifier is gemini-3.8-flash-lite-tts.


Sources 6
Provider

About Google DeepMind