What is Gemini 3.8 Flash-Lite TTS?
Gemini 3.8 Flash-Lite TTS is a text-to-speech model from Google. Its canonical Gemini API model identifier is gemini-3.8-flash-lite-tts. Unlike a general-purpose Gemini language model, it is focused on turning supplied text into spoken audio.
Google positions the model as the workhorse option in the Gemini 3.8 TTS family. The Flash-Lite designation indicates a focus on high throughput, low latency, and cost efficiency. It uses the same API schema and structured prompting approach as Gemini 3.8 Flash TTS, but is intended for workloads where economical, scalable speech generation matters more than maximum acoustic fidelity or highly nuanced studio-style performance.
The model was released as generally available on September 22, 2026. The information and introductory pricing described here reflect the supplied Google documentation and pricing details for September 2026.
How the model handles input and output
Flash-Lite TTS accepts text input and produces audio output. The supplied text is treated as the transcript to be spoken. Instructions about sustained characteristics such as tone, pace, whispering, or delivery style should be provided through structured speech_metadata rather than inserted into the transcript as ordinary stage directions.
The model supports both single-speaker and multi-speaker generation. For a multi-speaker request, speaker labels should be specified in speech metadata for every turn. This makes it possible to generate dialogue, narrated conversations, or other content that requires more than one voice.
- Input: Text
- Output: Spoken audio
- Speakers: Single-speaker and multi-speaker synthesis
- Languages: More than 100
- Default unary format: WAV
- Other audio formats: Raw PCM, mu-law, and A-law are available as options
- Input limit: 8,192 tokens
- Output serving limit: 16,384 tokens
Unary requests return WAV audio by default. Applications migrating from older Gemini TTS preview models should take care not to add a second WAV header to the returned bytes.
Voice and speech capabilities
The model includes several capabilities beyond basic text-to-speech. It can use prebuilt voices, the Extended Voice Library, custom Voice design personas, and Voice replication. Voice design allows an application to describe the kind of voice it wants, while Voice replication is intended for workflows that need a voice reproduced from an appropriate source or configuration.
Structured speech metadata can describe speakers and delivery style. The model also supports inline vocal events, including laughter, sighs, coughs, breaths, and short pauses. These events can be represented with tags such as <laugh>, <sigh>, and <short pause>. They are useful when a script needs some conversational texture rather than a completely uniform read.
These controls do not turn Flash-Lite TTS into a general audio-production or music-generation system. Its purpose remains speech synthesis from text.
Where Gemini 3.8 Flash-Lite TTS fits best
Flash-Lite TTS is most useful when an application needs speech generated frequently, with predictable API-based processing and controlled cost. Suitable examples include:
- High-volume production of narrated or read-aloud content
- Accessibility features that read articles, documents, or interface text aloud
- Low-latency responses in conversational voice-agent cascades
- Everyday assistant responses where speech must be generated quickly
- Multi-speaker dialogue and scripted conversational content
- Voice replication workflows that require scalable synthesis
- Applications that need speech in a broad range of languages
It is particularly attractive for pipelines in which a separate system handles conversation logic, retrieval, or reasoning and Flash-Lite TTS produces the final spoken response. The model itself should not be selected as the reasoning component of that pipeline: it is a speech-generation model, not a general-purpose chat or reasoning model.
Pricing, limits, and availability
Google lists the following standard paid pricing for Gemini 3.8 Flash-Lite TTS through December 31, 2026:
| Usage type | Price |
|---|---|
| Text input | $0.50 per 1 million text tokens |
| Audio output | $6.00 per 1 million audio tokens |
| Standard output equivalent | Approximately $0.0015 per 10 seconds of audio |
Beginning January 1, 2027, the listed standard prices are scheduled to increase to $1.00 per 1 million text input tokens and $12.00 per 1 million audio output tokens. Batch and Flex input is listed at $0.25 per 1 million tokens through December 31, 2026, increasing to $0.50 from January 1, 2027.
Audio billing uses 25 tokens per second of audio. The exact amount paid for a request therefore depends on the generated duration as well as the amount of text submitted. The introductory prices are date-limited, so production cost estimates should account for the scheduled 2027 increase rather than assuming the temporary rates will continue.
The model has an 8,192-token input limit and a 16,384-token output serving limit. Google lists caching, the Batch API, Flex inference, and Priority inference as supported. It is generally available rather than a preview-only model.
Supported and unsupported API features
Flash-Lite TTS supports speech generation features, but it does not expose the broader tool and output capabilities associated with some general-purpose Gemini models. The supplied specifications indicate that the following are not supported for this exact model:
- Function calling
- Code execution
- File search
- Live API
- Search grounding
- Google Maps grounding
- Structured outputs
- Thinking
- URL context
- Image generation
It also does not accept audio, image, or video input according to the supplied model specifications. Its multimodal capability is therefore limited to producing audio from text, rather than accepting arbitrary media for analysis or participating in audio-to-audio conversations.
Speed, cost, and capability trade-offs
The central choice behind Flash-Lite TTS is a production trade-off. Its design favors speed and cost control over the most detailed expressive performance. For a service generating thousands or millions of spoken responses, lower per-request cost and low latency can matter more than subtle improvements in acting, acoustic detail, or dialect-specific performance.
A premium speech option may be more appropriate when the primary requirement is highly expressive narration, maximum acoustic fidelity, extensive dialect coverage, or studio-quality voice performance. The supplied research specifically positions Gemini 3.8 Flash TTS as the more suitable sibling when those qualities are more important. Flash-Lite TTS is the better fit when the workload is repetitive, scalable, and price-sensitive.
The model is also a poor fit for tasks that require reasoning, coding, transcription, image generation, video generation, or structured JSON responses. Those functions should be handled by a different model or service before the resulting text is passed to Flash-Lite TTS.
Implementation guidance
Applications should separate transcript content from speech instructions. For example, the words intended to be spoken should remain in the text transcript, while persistent directions about pace, tone, or delivery should be represented through speech metadata. In multi-speaker requests, every turn should identify its speaker through the same structured mechanism.
Inline event tags are appropriate for momentary vocal actions. A script can include tags for a laugh, sigh, breath, cough, or short pause where supported. These should be used for local events rather than as a replacement for sustained style instructions.
Because unary responses are WAV by default, audio-handling code should preserve the returned format correctly. Adding another WAV header can corrupt the output. If an application needs raw PCM, mu-law, or A-law, it should request and process the selected format consistently throughout the pipeline.
When should you choose Gemini 3.8 Flash-Lite TTS?
Choose Gemini 3.8 Flash-Lite TTS when you need text-to-speech that is fast, comparatively inexpensive, and suitable for repeated production use. It is a strong candidate for read-aloud applications, high-volume narration, voice-agent cascades, everyday assistant speech, and voice workflows that benefit from single- or multi-speaker generation.
Choose another option when the project depends on premium studio narration, the broadest language or dialect coverage, maximum expressive control, audio-to-audio interaction, transcription, reasoning, or tool use. In particular, the flagship Gemini 3.8 Flash TTS model may be more appropriate when expressive fidelity is more important than the Flash-Lite model's throughput and cost advantages.
Overall assessment
Gemini 3.8 Flash-Lite TTS is a focused speech model rather than a general AI assistant. Its value comes from combining text-to-speech generation, multi-speaker support, voice design, voice replication, vocal events, broad language coverage, and production-oriented pricing in a model designed for scale.
Its limitations are equally important: it does not reason, call tools, analyze uploaded media, return structured JSON, or provide Live API audio-to-audio interaction. For organizations that need a dependable speech layer after another system has produced the text, those restrictions are usually a clear boundary rather than a disadvantage. The model's practical appeal is greatest when fast and economical speech generation is the requirement.

