Voxtral TTS

Voxtral TTS

by Mistral AI · GA; currently available through the Mistral API and Mistral Studio

Voxtral TTS is Mistral AI’s specialized speech-generation model for expressive audio in nine languages. It supports short-reference voice adaptation, cross-lingual voice transfer, streaming, saved voices, and MP3, WAV, PCM, FLAC, and Opus output. The model costs $0.016 per 1,000 output characters, generates up to two minutes natively, and is licensed CC BY-NC 4.0.

Speech Reasoning Coding
Voxtral TTS turns written text into natural-sounding speech with support for nine languages, reference-voice adaptation, cross-lingual voice transfer, streaming, and several audio formats. Available through the Mistral API and Mistral Studio, it is designed for multilingual voice agents, narration, accessibility, and other applications that need controllable speech rather than text responses.
Outputs

What Voxtral TTS can produce

Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Voxtral TTS
Model type Text To Speech
Release date 2026-03-23
Status GA; currently available through the Mistral API and Mistral Studio
Knowledge cutoff notes

A conventional textual knowledge cutoff is not published for this speech-generation model. It is designed for text-conditioned audio generation rather than factual question answering.

Model notes

Canonical API model ID is voxtral-mini-tts-2603. Voxtral TTS is a 4B-parameter text-to-speech model consisting of a 3.4B transformer decoder, a 390M flow-matching acoustic transformer, and a 300M neural audio codec. It supports English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic. The model supports zero-shot voice adaptation from short reference audio, including cross-lingual voice transfer, and can generate up to two minutes of audio natively; the API supports longer generations through interleaving. Mistral launch material describes references as short as 3 seconds, while the architecture documentation describes a 5-to-25-second voice prompt. Available response formats include MP3, WAV, PCM, FLAC, and Opus. The model is licensed CC BY-NC 4.0, so the license is not an unrestricted commercial open-source model license. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input $0 per 1 million input characters
Output $16 per 1 million output characters
Model guide

Voxtral TTS: Mistral AI’s Multilingual Voice Generation and Cloning Model

Voxtral TTS is Mistral AI’s 4-billion-parameter text-to-speech model for generating expressive speech in nine languages, adapting a voice from short reference audio, transferring voices across languages, and delivering low-latency audio through streaming API responses.

What is Voxtral TTS?

Voxtral TTS is Mistral AI’s text-to-speech model for producing spoken audio from written input. Its canonical API model ID is voxtral-mini-tts-2603. Mistral released it on March 23, 2026, and the model is currently generally available through the Mistral API and Mistral Studio.

The model is aimed at speech generation rather than general-purpose reasoning. It accepts text and returns audio, making it suitable for applications such as voice assistants, multilingual narration, conversational agents, accessibility tools, and customized spoken content. It is not a transcription model, a general chat model, an image model, or a video-generation system.

Mistral describes Voxtral TTS as a 4-billion-parameter model. Its architecture combines a 3.4-billion-parameter transformer decoder, a 390-million-parameter flow-matching acoustic transformer, and a 300-million-parameter neural audio codec. These components handle text conditioning, acoustic generation, and conversion into usable audio.

Languages and voice generation

Voxtral TTS supports English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic. That language coverage makes it more relevant to multilingual products than a speech system limited to one or two major languages.

The model’s most distinctive capability is voice adaptation. An application can provide reference audio, and Voxtral TTS can generate speech that adapts to the referenced voice without requiring a transcript of that recording. Mistral’s launch material describes reference clips as short as three seconds, while the architecture documentation describes a voice prompt of approximately five to 25 seconds. Because the published guidance differs, implementations should validate the minimum and recommended reference duration against the current API behavior before relying on a specific threshold.

Voice adaptation also supports cross-lingual transfer. For example, a reference voice in one supported language can be used to generate speech in another supported language. This can help organizations maintain a consistent narrator or assistant identity across localized versions of an application, although voice quality and pronunciation should be evaluated separately for each language.

Inputs, outputs, and documented limits

The documented input modality is text, optionally accompanied by reference audio for voice adaptation. The direct model output is audio rather than text, images, video, or structured data. Voxtral TTS supports streaming, which allows audio to begin arriving before a complete generation has finished and is particularly useful for interactive voice applications.

Mistral states that the model can generate up to two minutes of audio natively. The API supports longer generations through interleaving, according to the supplied model notes. Interleaving should not be interpreted as an unlimited output mode: longer content may require application-level segmentation and assembly, and developers should test continuity, latency, and any operational limits for their workload.

A conventional context-window or maximum-token specification is not published for this speech-generation model. The relevant practical limits are therefore the accepted text size, generation duration, reference-audio behavior, and API constraints documented for the account and endpoint. Applications should avoid assuming that a language-model context window applies to Voxtral TTS.

Available response formats include MP3, WAV, PCM, FLAC, and Opus. This range covers common downloadable, editing, raw-audio, and streaming-oriented workflows. The appropriate format depends on the application: compressed formats can reduce transfer size, while PCM or WAV may be more convenient for processing and production pipelines.

Pricing and API availability

The published price is $0.016 per 1,000 output characters, equivalent to $16 per 1 million output characters. Input text is priced at $0 per 1 million input characters according to the supplied model data. Since billing is based on output characters rather than audio seconds, costs will vary with the amount of generated text, language, and the number of requested generations.

Voxtral TTS is available through the Mistral API and Mistral Studio. The API documentation covers standard speech generation, saved voices, reference-audio prompts, streaming, and response-format selection. Developers should treat the model ID, request schema, authentication requirements, quotas, and operational limits in the current documentation as authoritative because these details can change independently of the model’s underlying capabilities.

Main strengths and trade-offs

  • Multilingual coverage: Nine supported languages make the model useful for products that need one speech system across multiple markets.
  • Voice adaptation: Short reference audio can be used to create a recognizable voice without a supplied transcript, and the model supports cross-lingual voice transfer.
  • Expressive speech: Mistral positions the model for natural and emotionally expressive output rather than purely mechanical reading.
  • Streaming: Streaming responses are useful when perceived latency matters, such as in voice agents and interactive interfaces.
  • Format flexibility: MP3, WAV, PCM, FLAC, and Opus support different delivery and processing requirements.
  • Cost profile: The published character-based price is comparatively straightforward to estimate for text-heavy workloads, while free input pricing avoids a separate charge for submitted text.

The trade-offs are equally important. Voxtral TTS is specialized for speech generation and does not provide general reasoning, coding, web search, image generation, transcription, embeddings, or function-calling capabilities in the supplied specification. Its voice adaptation license is also a significant consideration: the model is licensed under CC BY-NC 4.0, so it should not be treated as an unrestricted commercial open-source model-weight license.

The model’s supported languages do not guarantee identical pronunciation, prosody, or voice similarity in every language. Reference-voice use also creates product, consent, and identity-management responsibilities. Teams should obtain appropriate permission for recorded voices and test generated speech before deploying it in customer-facing or high-stakes settings.

Speed, cost, and capability positioning

The supplied comparative editorial assessment rates Voxtral TTS highly for speed and cost, with a speed score of 9 out of 10 and a cost score of 8 out of 10. These are editorial estimates, not provider-published benchmarks, and they should not be read as measured latency or a guaranteed price advantage in every workload.

In practical terms, Voxtral TTS prioritizes focused speech generation over the broader capabilities of a general-purpose language model. A specialized text-to-speech model can be a better fit than asking a conversational model to manage speech as one part of a larger workflow when the main requirement is consistent, multilingual audio output. Streaming and saved voices can also reduce application complexity for repeated voice experiences.

On the other hand, a broader system may be more appropriate when the application must reason over user requests, call tools, generate code, search the web, or produce multiple media types. Such a system could generate the text first and then pass it to a speech model, whereas Voxtral TTS itself should be selected for the speech-generation stage.

Best use cases

Voxtral TTS is a strong candidate for applications where spoken output is the primary deliverable:

  • Multilingual voice agents: Generate spoken responses in supported languages and stream them to users with low perceived delay.
  • Localized narration: Produce versions of educational, instructional, or media content across multiple languages while retaining a consistent adapted voice.
  • Accessibility features: Add speech output to applications for users who benefit from audio interaction or spoken content.
  • Interactive characters and assistants: Use reference audio to establish a recognizable voice identity for an assistant or fictional character.
  • Audio prototyping: Quickly test narration, dialogue, and voice-interface concepts without recording every variation manually.
  • Content pipelines: Select among compressed and lossless formats depending on whether the output is for delivery, editing, or downstream audio processing.

For these uses, teams should test long-form continuity, pronunciation of names and specialist terms, emotional delivery, reference-voice similarity, and the latency of streaming responses. A short demonstration can sound good while still revealing problems at production scale.

When to choose Voxtral TTS

Choose Voxtral TTS when you need a Mistral-hosted speech model that combines multilingual generation, expressive delivery, short-reference voice adaptation, cross-lingual voice transfer, streaming, and multiple audio formats. It is especially attractive when one product must serve several of the nine supported languages and when voice identity matters as much as the words being spoken.

Consider another option when you need commercial terms broader than CC BY-NC 4.0, a published enterprise voice-governance framework tailored to your requirements, transcription rather than speech generation, music or sound-effect creation, or a general model that can reason and take actions directly. A separate reasoning or orchestration model may be needed before Voxtral TTS if the application must interpret complex requests, retrieve information, or decide which tools to call.

Voxtral TTS is therefore best understood as a focused audio-generation component, not as a complete conversational agent. Its value comes from turning prepared text into adaptable, multilingual speech efficiently; the surrounding application remains responsible for dialogue logic, content safety, permissions for reference voices, and quality control.

Bottom line

Voxtral TTS gives developers a dedicated route from text to expressive multilingual audio. Its combination of nine-language support, zero-shot voice adaptation, cross-lingual transfer, streaming, and broad output-format support distinguishes it from a basic single-voice reader. The main limitations are its specialized scope, unavailable conventional context and token limits, native two-minute generation limit, and CC BY-NC 4.0 license.

For voice agents, localized narration, accessibility, and applications that need a reusable voice identity, it is a focused option worth evaluating. For reasoning, coding, transcription, unrestricted commercial model licensing, or multimodal application control, it should be paired with or replaced by a more suitable system.


Answers to Frequently Asked Questions

What are the main limitations of Voxtral TTS?
Voxtral TTS is specialized for speech generation and does not provide general reasoning, transcription, coding, web search, image generation, or function calling. It can generate up to two minutes of audio natively, and its CC BY-NC 4.0 license may not suit unrestricted commercial use.
How much does Voxtral TTS cost and which audio formats does it support?
The published price is $0.016 per 1,000 output characters, or $16 per 1 million output characters, while input text is listed at $0 per 1 million characters. Supported output formats include MP3, WAV, PCM, FLAC, and Opus.
Can Voxtral TTS clone or adapt a person’s voice?
Yes. Voxtral TTS can adapt generated speech to a reference voice without requiring a transcript of the recording. Mistral describes support for short reference clips and cross-lingual voice transfer, but developers should verify current duration requirements and obtain permission to use recorded voices.
What is Voxtral TTS and what is it used for?
Voxtral TTS is Mistral AI’s text-to-speech model for converting written text into spoken audio. It is designed for multilingual voice agents, narration, accessibility tools, conversational applications, and customized spoken content.
Which languages does Voxtral TTS support?
Voxtral TTS supports English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic.


Sources 4
Provider

About Mistral AI