Magpie TTS

Magpie TTS Multilingual

by NVIDIA AI · Current

NVIDIA Magpie TTS Multilingual is a 357M-parameter text-to-speech model for generating expressive spoken audio in 12 languages. It supports built-in voices, selected emotional styles, cross-language speaker use, text normalization, streaming, offline inference, NeMo Speech customization, and NVIDIA NIM deployment. Standard generations can produce up to 20 seconds of speech, while long-form generation is beta. The current checkpoint does not support zero-shot voice cloning.

Speech Reasoning Coding
NVIDIA Magpie TTS Multilingual is designed for applications that need consistent, natural-sounding speech across multiple languages. The current v2607 checkpoint supports 12 languages, multiple built-in speakers, selected emotional voice variants, streaming and offline inference, and standard generations of up to 20 seconds. It is a focused speech-synthesis model rather than a general-purpose language model, and it does not provide zero-shot voice cloning.
Outputs

What Magpie TTS Multilingual can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Magpie TTS
Model type Other
Release date 2026-07-21
Status Current
Knowledge cutoff notes

Knowledge-cutoff metadata is not applicable to this text-to-speech model. It synthesizes speech from supplied text and does not function as a knowledge-grounded language model.

Model notes

The current v2607 checkpoint is documented as a 357M-parameter multilingual text-to-speech model and supports Arabic, Chinese, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Spanish and Vietnamese. It produces spoken audio rather than text. NVIDIA NIM supports streaming and offline inference with locale-specific voices and batch-size profiles. Standard inference supports up to 20 seconds of speech per generation; long-form generation uses a sliding-window mode and is documented as beta. The current checkpoint removed zero-shot voice cloning for security reasons. NVIDIA documentation supports training and customization through NeMo Speech, including preference optimization and custom pronunciation resources. No public per-unit production price for this exact model was verified; the NVIDIA Build page exposes a free development API that may be rate limited.

Model guide

NVIDIA Magpie TTS Multilingual: Expressive Speech in 12 Languages

NVIDIA Magpie TTS Multilingual is a 357M-parameter text-to-speech model that converts text into expressive spoken audio in 12 languages. It offers built-in voices, streaming and offline inference, text normalization, cross-language voice use, and deployment through NeMo Speech, NVIDIA NIM, or NVIDIA's hosted development API.

What is Magpie TTS Multilingual?

NVIDIA Magpie TTS Multilingual is a 357-million-parameter text-to-speech model from NVIDIA. It takes written text as input and produces spoken audio, allowing software to add narration, voice interfaces, accessibility features, dubbing, or interactive characters without maintaining separate speech models for every supported language.

The current documented checkpoint is v2607. NVIDIA distributes it through its model ecosystem, including the NVIDIA Hugging Face organization, NeMo Speech workflows, NVIDIA NIM deployments, and a hosted development endpoint on NVIDIA Build. Its role in NVIDIA's catalog is specialized: unlike a conversational language model, Magpie TTS is intended to perform speech generation after an application has already determined what should be said.

For example, a voice assistant could use a separate language model to write a response and then pass that response to Magpie TTS Multilingual for playback. The TTS model itself does not provide general reasoning, web search, text generation, or knowledge retrieval.

Supported languages and voice options

Magpie TTS Multilingual supports Arabic, Chinese, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Spanish, and Vietnamese. This makes it useful for products that need to deliver the same type of spoken experience across several markets.

NVIDIA's NIM deployment exposes locale-specific voices and multiple built-in speakers. Documented speaker names include Aria, Jason, Leo, Sofia, Mia, Ray, Diego, Pascal, Isabela, Louise, HouZhen, Long, Siwei, and Phung. The exact voice availability depends on the language and deployment version rather than every speaker being available for every locale.

The model can preserve a speaker identity across supported languages. It also supports selected emotional voice variants and can combine a voice associated with one locale with text in another supported language to create accented speech. These features can help when an application needs one recognizable character or narrator across a multilingual experience.

How the model generates speech

Magpie TTS Multilingual uses an encoder-decoder transformer architecture. In practical terms, the encoder processes the supplied text and the decoder autoregressively predicts discrete audio codec tokens, which are compact representations of speech. A pretrained neural audio codec then converts those tokens into waveform audio that can be played or saved.

The model uses a two-stage decoding design. Frame stacking allows the main decoder to process groups of consecutive audio frames, while a lightweight local transformer refines frame-level predictions. This design separates broader sequence generation from detailed frame-level speech reconstruction.

NVIDIA documents attention priors to improve alignment between text and speech and to reduce the risk of skipped, repeated, or otherwise incorrect content on difficult inputs. Classifier-free guidance provides another control mechanism for balancing adherence to the conditioning information with generation diversity.

The current model documentation also describes preference optimization and IPA grapheme-to-phoneme support. IPA, or the International Phonetic Alphabet, gives developers a way to represent pronunciation more explicitly. Custom pronunciation resources can therefore help with names, specialist terms, or code-switching where ordinary text normalization might not produce the desired pronunciation.

Inputs, outputs, and generation limits

The model accepts text and produces spoken audio. It does not accept image, audio, or video input according to the supplied model specifications, and it does not return text, images, video, music, embeddings, or structured data as its primary output.

CapabilityDocumented behavior
InputText
OutputSpoken audio
Languages12 documented languages
Standard generation lengthUp to 20 seconds of speech per generation
Long-form speechSliding-window mode, documented as beta
StreamingSupported through NVIDIA TTS NIM
Offline inferenceSupported through NVIDIA TTS NIM and local NeMo Speech workflows
Voice cloningZero-shot voice cloning is not included in the current checkpoint

Text normalization is required. Numbers, abbreviations, punctuation, and other written forms may need to be converted into speech-friendly text before synthesis. For long passages, NVIDIA recommends clear punctuation and sentence boundaries. The sliding-window approach can extend generation beyond the standard single-generation limit, but its beta status means applications should test continuity, pronunciation, and transitions carefully.

Deployment through NeMo, NIM, and hosted access

There are three main ways to work with Magpie TTS Multilingual in the supplied NVIDIA documentation.

  • NeMo Speech: A local or self-managed workflow for inference, training, and customization within NVIDIA's speech framework.
  • NVIDIA NIM: A deployable inference service that supports streaming and offline operation, locale-specific voices, and batch-size profiles.
  • NVIDIA Build: A hosted development endpoint that provides access for experimentation. The supplied research describes this development API as free, although it may be rate limited.

NIM deployment requires a compatible NVIDIA GPU with Compute Capability 8.0 or higher. That requirement makes the model more suitable for NVIDIA-equipped servers, workstations, and controlled application environments than for arbitrary consumer hardware. Hardware, driver, container, licensing, and operational requirements should be checked separately for a particular deployment.

No public per-unit production price for this exact model was verified in the supplied research. The free hosted development access should not be interpreted as a confirmed unlimited or production pricing plan. Organizations planning substantial production traffic should verify current NVIDIA terms directly.

Main strengths and trade-offs

The clearest strength of Magpie TTS Multilingual is its combination of language coverage and a shared voice experience. A product can support Arabic, Chinese, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Spanish, and Vietnamese without selecting a completely different TTS system for each language. Cross-language speaker support may also help maintain a consistent brand, narrator, or fictional character.

Streaming and offline modes provide different integration choices. Streaming can reduce the time before speech starts in an interactive assistant, while offline inference is useful when an application needs local processing, controlled infrastructure, or batch generation. Text normalization, pronunciation resources, and IPA support give developers more control than a simple text-in/audio-out interface.

However, the model is deliberately narrower than a general AI system. It does not reason about a request, write an answer, call tools, browse the web, or analyze files. It should normally be paired with an application layer or another model when the product needs dialogue management, content generation, retrieval, or actions.

The 12-language list is also a boundary, not merely a recommended starting point. Text outside the documented language set is not supported by the supplied research. The standard 20-second generation limit requires additional handling for longer speech, and beta long-form generation may require testing for natural continuity. Finally, the current checkpoint does not provide zero-shot voice cloning, so it is not the appropriate choice when users must create speech from an arbitrary reference recording.

Capability profile

Magpie TTS Multilingual should be evaluated as a speech-generation component rather than by the criteria used for large language models.

  • Reasoning: Not a reasoning model. It synthesizes supplied text and does not independently solve problems or make plans.
  • Coding: Not intended for code generation or code understanding.
  • Tool use: No native tool or function-calling capability is documented.
  • Structured output: Not applicable as a primary model output; the output is audio.
  • Streaming: Supported through NVIDIA TTS NIM.
  • Customization: NVIDIA documents training and customization through NeMo Speech, including preference optimization and custom pronunciation resources.
  • Speed and cost: The model is optimized for dedicated NVIDIA inference workflows and is rated editorially as speed-oriented, but the supplied research does not provide a benchmark or verified per-character, per-second, or per-request production price.

These distinctions matter when comparing Magpie TTS with another option. A general language model may be better for generating the content to speak, while a voice-cloning system may be better when an application requires a custom voice derived from a recording. A broader speech platform may also be more appropriate if the project needs unsupported languages or audio understanding in addition to speech synthesis.

Best use cases

  • Multilingual voice agents: Generate spoken replies after a dialogue system has created the response text.
  • Accessibility and screen reading: Convert interface content or written documents into speech in supported languages.
  • Audiobooks and narration: Produce consistent narration for short sections, previews, or multilingual editions.
  • Dubbing and localization: Create localized speech tracks while retaining a selected built-in speaker identity.
  • Digital humans and interactive characters: Give a character a repeatable voice across several supported locales.
  • Interactive media: Use streaming inference for applications where speech should begin before the complete response has finished processing.

When to choose Magpie TTS Multilingual

Choose Magpie TTS Multilingual when the central requirement is expressive, multilingual text-to-speech and the supported NVIDIA deployment environment fits your infrastructure. It is particularly suitable when a project values 12-language coverage, built-in voices, consistent speaker identity, streaming, or the ability to customize pronunciation through NeMo Speech.

Choose another type of model when the primary task is open-ended conversation, reasoning, coding, image or video generation, speech recognition, or voice creation from an arbitrary reference speaker. Magpie TTS can be one component in those systems, but it is not a replacement for them.

It is also worth considering another speech solution when the required language is outside NVIDIA's documented set, when production pricing must be predictable before deployment, or when long-form generation and voice cloning are core requirements. For supported languages and NVIDIA GPU infrastructure, however, Magpie TTS Multilingual offers a focused path from prepared text to expressive spoken audio.


Answers to Frequently Asked Questions

What are the main generation limits of NVIDIA Magpie TTS Multilingual?
Standard generation supports up to 20 seconds of speech per generation. Longer passages can use the documented beta sliding-window mode, which should be tested for continuity and pronunciation. Text normalization is required, and custom pronunciation resources or IPA input can help with names and specialist terms.
How can developers deploy NVIDIA Magpie TTS Multilingual?
Developers can use local or self-managed NeMo Speech workflows, deploy the model through NVIDIA TTS NIM, or access it through the hosted NVIDIA Build development endpoint. NIM supports streaming and offline inference, while compatible deployments require an NVIDIA GPU with Compute Capability 8.0 or higher.
Can NVIDIA Magpie TTS Multilingual clone any speaker's voice?
No. The current checkpoint does not include zero-shot voice cloning from an arbitrary reference recording. It provides built-in speakers, supports selected emotional voice variants, and can preserve a speaker identity across supported languages.
What is NVIDIA Magpie TTS Multilingual?
NVIDIA Magpie TTS Multilingual is a 357-million-parameter text-to-speech model that converts written text into spoken audio. It is designed for narration, voice interfaces, accessibility, dubbing, and interactive characters rather than reasoning, text generation, or web search.
Which languages does NVIDIA Magpie TTS Multilingual support?
The model supports 12 languages: Arabic, Chinese, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Spanish, and Vietnamese.


Sources 5
Provider

About NVIDIA AI