MAI-Voice

MAI-Voice-2

by Microsoft Copilot · Public preview

Microsoft MAI-Voice-2 is a high-fidelity multilingual text-to-speech model for expressive narration, audiobooks, podcasts, voice-over, accessibility, and branded audio. It supports more than 15 languages and 18 locales, granular emotional delivery, long-form speaker consistency, and gated zero-shot voice prompting from short reference clips. The listed price is $22 per 1 million characters.

Speech Reasoning Coding
MAI-Voice-2 is Microsoft's highest-fidelity MAI-Voice model for natural and expressive speech generation. Available through Microsoft Foundry and Azure Speech, it supports more than 15 languages and 18 locales, with controls for emotion, delivery style, and speaker identity. The model is aimed at quality-sensitive audio production rather than the lowest possible response latency.
Outputs

What MAI-Voice-2 can produce

Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family MAI-Voice
Model type Other
Release date 2026-06-02
Status Public preview
Knowledge cutoff notes

Microsoft's public model and Azure Speech documentation reviewed for this record does not specify a knowledge cutoff. The model is a text-to-speech system, so a language-model knowledge cutoff is not a directly applicable published specification.

Model notes

MAI-Voice-2 is a speech-generation model rather than a general-purpose language model. It accepts text and can use a short reference audio clip for gated zero-shot voice prompting. Microsoft documents support for more than 15 languages and 18 locales, granular emotion control, long-form speaker consistency, and approximately one second of model inference for 45 seconds of generated audio. The listed price is character-based rather than token-based. Personal voice and voice-prompting access require Microsoft approval and consent safeguards; only authorized, licensed voices may be synthesized in production. Editorial scores are comparative estimates and are not vendor benchmarks.

Cost

Model pricing

Input $22 per 1 million characters
Output $22 per 1 million characters
Model guide

Microsoft MAI-Voice-2 for Expressive Multilingual Narration and Voice Creation

Microsoft MAI-Voice-2 is a high-fidelity text-to-speech model for expressive multilingual speech, long-form narration, voice-over, branded audio, and accessibility experiences. It supports granular emotional delivery, maintains speaker consistency over extended content, and offers gated zero-shot voice prompting from a short reference recording.

What is Microsoft MAI-Voice-2?

MAI-Voice-2 is Microsoft's expressive text-to-speech model. It converts written text into spoken audio with emphasis on natural pacing, tone, pitch, stress, emotional delivery, and consistent speaker identity. In practical terms, it is designed for producing finished or near-finished narration rather than simply reading text in a neutral computer voice.

The model is available through Microsoft Foundry and Azure Speech. Microsoft positions it for long-form narration, audiobooks, podcasts, educational content, documentaries, voice-over, accessibility experiences, assistants, and branded audio. Its central trade-off is straightforward: MAI-Voice-2 prioritizes fidelity and expressive control over the lowest possible latency.

Where MAI-Voice-2 fits in Microsoft's voice model lineup

MAI-Voice-2 belongs to Microsoft's MAI-Voice family and is positioned as the higher-fidelity option for expressive speech generation. The supplied Microsoft material contrasts it with MAI-Voice-2-Flash, which is more appropriate when rapid response times and high-volume interactive conversations are the main priorities.

This positioning makes MAI-Voice-2 closer to a production narration and voice-over system than to a lightweight conversational voice endpoint. A content producer may accept its roughly one-second inference time for 45 seconds of generated audio because the output needs to sound consistent and emotionally appropriate across a long recording. A real-time assistant or call-center system may instead prefer a faster model, even if that means accepting different quality or control trade-offs.

Core capabilities

  • Expressive speech generation: The model generates speech with natural pacing, pitch, stress, tone, and emotional range.
  • Granular delivery control: Supported styles include excitement, sadness, anger, fear, whispering, shouting, soft voice, and surprise. The documented style set also includes angry, confused, determined, disgusted, embarrassed, fearful, happy, hopeful, jealous, joyful, regretful, relieved, and other expressive variations.
  • Long-form consistency: Microsoft describes the model as maintaining speaker identity and stable voice quality across audiobooks, lectures, podcasts, courses, and other extended material.
  • Multilingual synthesis: It supports more than 15 languages and 18 locales, including English, German, French, Italian, Hindi, Spanish, Portuguese, Korean, Chinese, Turkish, Russian, Thai, Dutch, Romanian, and Hungarian.
  • Code-switching: Microsoft reports support for selected mixed-language scenarios, including Hindi-English and Spanish-English.
  • Voice prompting: A short reference recording can be used to guide the generated voice without conventional model fine-tuning. This capability is gated and subject to consent and authorization requirements.

Inputs, outputs, and unavailable model functions

The primary input is text. A short audio reference clip can also be used for the gated voice-prompting workflow. Microsoft documentation recommends a reference clip of approximately 5 to 60 seconds for the custom voice process.

The direct output is generated speech audio. MAI-Voice-2 is not a general-purpose language model and does not produce text, images, video, embeddings, or music as its primary model output. It also does not provide general reasoning, coding, web search, or tool/function execution. Applications that need script writing, dialogue generation, content planning, or other language-model tasks must supply those functions separately through an upstream model or application workflow.

No published context-window limit or maximum output-token limit is provided in the supplied documentation. Those language-model metrics are also not a natural description of a speech-generation system. The relevant production constraints are instead the amount of text submitted, audio duration, supported languages and voices, service limits, and the availability of the required Azure or Foundry resource.

Languages, locales, and voice styles

Microsoft describes MAI-Voice-2 as supporting 15 or more languages and 18 locales. Documented coverage includes United States and Australian English; German; French; Italian; Hindi; European and Mexican Spanish; Brazilian and European Portuguese; Korean; Simplified Chinese; Turkish; Russian; Thai; Dutch; Romanian; and Hungarian.

Locale support does not mean every voice has identical expressive behavior. Available voices and styles vary by language and locale. Before committing to a production voice, teams should test pronunciation, names, abbreviations, code-switching, emotional delivery, and long-form consistency in the target language. This is particularly important for branded narration, where a technically supported locale may still require editorial review for accent and pronunciation quality.

Voice prompting, consent, and safety controls

MAI-Voice-2 can use a short reference clip to create speech that resembles the referenced speaker. This is commonly described as zero-shot voice prompting because the application does not need to fine-tune the model for each new voice. The feature can be useful for approved personal voices, characters, narration styles, or brand-controlled audio identities.

However, the capability is not presented as unrestricted voice cloning. Microsoft states that personal voice and voice-prompting access is gated, and that consent safeguards are required. Only authorized and licensed voices may be synthesized for production use. Organizations should therefore plan for speaker authorization, consent records, content review, and access controls rather than treating a reference recording as sufficient permission by itself.

Performance and pricing

Microsoft reports approximately one second of model inference time for generating 45 seconds of audio. This is a useful production-oriented measure: it indicates that the system can generate long passages substantially faster than real-time playback, while still not being positioned as the lowest-latency model in the MAI-Voice family. Actual end-to-end time can also include request handling, queueing, audio processing, and application overhead.

The listed price is $22 per 1 million characters. This is character-based pricing, not token-based pricing. The supplied research does not provide a separate monthly subscription price, output-duration rate, context limit, or guaranteed service quota. Actual billing, regional availability, account requirements, and Azure service conditions may vary, so the displayed model price should be treated as the documented list price rather than a promise of a particular project total.

Main strengths and limitations

Strengths

  • Expressive control: Emotional and delivery styles make it more suitable for narration and characterful voice-over than a basic neutral TTS system.
  • Long-form suitability: Speaker consistency is important for audiobooks, courses, podcasts, and documentaries, where noticeable voice drift can make an otherwise usable recording difficult to edit.
  • Broad multilingual coverage: The documented language and locale range supports international content workflows.
  • Voice identity options: Gated reference-based prompting can support authorized voices without conventional per-speaker fine-tuning.
  • Production-oriented output: The model is designed around high-fidelity audio rather than general-purpose text generation.

Limitations

  • Not a general AI assistant: It does not reason over tasks, write scripts independently, generate code, browse the web, or call tools.
  • Not the best fit for minimum latency: Interactive applications requiring immediate turn-taking may benefit from MAI-Voice-2-Flash or another low-latency speech model.
  • Restricted voice prompting: Custom or personal voice capabilities require supported access, Microsoft approval, and consent safeguards.
  • Variable availability: Voices, languages, account requirements, and Azure resource access can depend on Microsoft's current documentation and regional service availability.
  • Requires verification: Pronunciation, emotional interpretation, multilingual delivery, and generated content should be reviewed before publication.

Best use cases

MAI-Voice-2 is a strong candidate when the audio itself is a central part of the user experience and quality matters more than instant response. Suitable applications include:

  • Audiobook and podcast narration
  • Documentary, course, lecture, and training content
  • Brand-controlled marketing voice-over
  • Accessibility narration and assistive voice experiences
  • Educational and entertainment characters
  • Long-form multilingual content with recurring speakers
  • Branded assistants where voice identity and expressive delivery matter more than minimum latency

For example, a course publisher could use an approved narrator voice across many lessons, while applying different emotional styles to introductions, warnings, examples, or story segments. A podcast workflow could generate draft narration in several supported locales and then route the audio through human review before release.

When to choose MAI-Voice-2

Choose MAI-Voice-2 when you need expressive, high-fidelity speech for content that may last minutes or hours, especially when consistent speaker identity and multilingual delivery are important. Its character-based pricing and reported generation speed can make it practical for batch production, provided the output is reviewed and the required voice permissions are in place.

Choose a faster speech option, including the MAI-Voice-2-Flash sibling identified in Microsoft's positioning, when the application depends on rapid conversational turn-taking, high-volume call-center interactions, or very low response time. Choose a separate language model when the system must reason, write or revise text, generate code, use tools, or manage a complex dialogue workflow. In many applications, MAI-Voice-2 is best used as the speech layer after another system has produced and approved the text.

Bottom line

Microsoft MAI-Voice-2 is a specialized expressive text-to-speech model for production-quality narration and branded audio. Its differentiators are emotional delivery control, multilingual support, long-form speaker consistency, and gated reference-based voice prompting. It is less suitable as an all-purpose conversational model or as the lowest-latency choice. The most important implementation checks are language and voice availability, consent for any reference voice, character-based cost, and audio review before public or commercial use.


Answers to Frequently Asked Questions

How does MAI-Voice-2 differ from MAI-Voice-2-Flash?
MAI-Voice-2 prioritizes expressive, high-fidelity speech, long-form consistency, and production narration. MAI-Voice-2-Flash is positioned as the better choice for rapid response times, high-volume interactive conversations, call centers, and applications requiring low-latency turn-taking.
How much does MAI-Voice-2 cost and how fast does it generate audio?
The listed price is $22 per 1 million characters. Microsoft reports approximately one second of model inference time to generate 45 seconds of audio, although total processing time may also include request handling, queueing, audio processing, and application overhead.
Can MAI-Voice-2 clone or replicate a person's voice?
MAI-Voice-2 supports gated voice prompting using a reference recording of approximately 5 to 60 seconds. This can create speech resembling the referenced speaker without conventional fine-tuning, but access is restricted and requires consent, authorization, and appropriate licensing. A recording alone does not establish permission to use a voice.
What is Microsoft MAI-Voice-2 used for?
Microsoft MAI-Voice-2 is an expressive text-to-speech model designed for production-quality narration, audiobooks, podcasts, documentaries, educational content, accessibility experiences, voice-over, assistants, and branded audio.
Which languages and voice styles does MAI-Voice-2 support?
MAI-Voice-2 supports more than 15 languages and 18 locales, including English, German, French, Italian, Hindi, Spanish, Portuguese, Korean, Chinese, Turkish, Russian, Thai, Dutch, Romanian, and Hungarian. Its expressive styles include excitement, sadness, anger, fear, whispering, shouting, soft voice, surprise, happiness, confusion, and other emotional variations. Available voices and styles vary by language and locale.


Sources 4
Provider

About Microsoft Copilot