MAI-Voice-2

MAI-Voice-2-Flash

by Microsoft Copilot · Public preview

Microsoft MAI-Voice-2-Flash is a public-preview text-to-speech model optimized for fast, expressive speech in real-time applications. It supports 15 languages and 18 locales, curated voices, SSML controls, streaming, and gated instant voice prompting through Azure Speech and Voice Live. Microsoft reports approximately 225 milliseconds of model-inference latency for 45 seconds of audio and lists pricing at $15 per one million characters. It is better suited to conversational systems than long-form narration requiring maximum fidelity.

Speech
MAI-Voice-2-Flash is Microsoft’s latency-focused model for turning text into natural-sounding speech. It is intended for interactive applications where a short delay matters more than maximum long-form audio fidelity, including voice assistants, customer-service automation, IVR systems, and multilingual conversational experiences.
Outputs

What MAI-Voice-2-Flash can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

10/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MAI-Voice-2
Model type Other
Release date 2026-07-23
Status Public preview
Knowledge cutoff notes

A model knowledge cutoff is not applicable or publicly specified for this text-to-speech model.

Model notes

MAI-Voice-2-Flash is a text-to-speech and audio-generation model available through Azure Speech in public preview. Microsoft describes approximately 225 ms of model-inference latency for generating 45 seconds of audio. It supports 15 languages and 18 locales, curated licensed voices, SSML-based expressive control, and gated instant voice prompting from a consented five- to 60-second audio clip without additional fine-tuning. Microsoft AI lists pricing at $15 per one million characters. Azure Speech pricing and availability may vary by region and service configuration. The model is optimized for latency-sensitive interaction; MAI-Voice-2 is the related model intended for higher-fidelity long-form generation.

Cost

Model pricing

Input $15 per 1 million characters
Output $15 per 1 million characters
Model guide

MAI-Voice-2-Flash: Microsoft’s Low-Latency Expressive Text-to-Speech Model

Microsoft MAI-Voice-2-Flash is a public-preview text-to-speech model designed for fast, expressive speech generation in real-time voice agents, assistants, call-center systems, and interactive voice response applications. It supports 15 languages and 18 locales, curated voices, SSML controls, and gated instant voice prompting through Azure Speech and Voice Live.

What is MAI-Voice-2-Flash?

MAI-Voice-2-Flash is a Microsoft text-to-speech model available through Azure Speech in public preview. It converts written text into generated speech and is designed for applications that need a quick spoken response rather than an extended production narration workflow.

The model belongs to Microsoft’s MAI-Voice-2 family. The “Flash” variant is positioned around responsiveness: Microsoft reports approximately 225 milliseconds of model-inference latency when generating 45 seconds of audio. That figure is a provider claim about model inference, not necessarily the total time an application user will experience. Network transfer, request handling, audio playback, and other service operations can add latency.

MAI-Voice-2-Flash is not a general-purpose language model. Its primary output is audio, while its input is text. It does not provide text, image, video, embedding, or general action output as its documented model function.

Where it fits in Microsoft’s lineup

MAI-Voice-2-Flash is the speed-oriented member of the MAI-Voice-2 family. Microsoft describes the related MAI-Voice-2 model as the better choice when the priority is higher-fidelity, longer-form generation, such as extended narration, podcasts, or audiobook-style content.

This makes the distinction practical rather than merely cosmetic. MAI-Voice-2-Flash is intended for short conversational turns and situations where the system should begin responding quickly. The related MAI-Voice-2 option may be more appropriate when consistent speaker quality and long-form fidelity matter more than immediate response time.

Speech capabilities

Expressive speech synthesis

MAI-Voice-2-Flash is designed to produce speech with natural rhythm, intonation, pacing, and emotional nuance. Supported SSML controls can be used to influence delivery, tone, and style at the voice or utterance level. SSML, or Speech Synthesis Markup Language, is a markup format that lets an application provide instructions about how text should be spoken.

These controls are useful for practical differences in delivery. For example, an application can structure a customer-service response so that a confirmation sounds clear and measured, while an assistant’s conversational reply uses a more natural pace. The exact controls and supported styles can vary by voice and locale.

Languages and locales

The model supports 15 languages and 18 locales. The documented language set includes English, Italian, French, German, Hindi, Spanish, Portuguese, Korean, Simplified Chinese, Turkish, Russian, Thai, Dutch, Romanian, and Hungarian.

Language support does not mean that every voice, accent, or expressive style is available in every locale. Applications should verify the specific voice and locale combination they plan to use, particularly when building a multilingual customer-facing service.

Curated voices and instant voice prompting

Applications can use Microsoft’s licensed, prebuilt voices. Voice identifiers use the MAI-Voice naming format, with examples including en-US-Harper:MAI-Voice-2-Flash and en-US-Ethan:MAI-Voice-2-Flash.

The model also supports gated instant voice prompting. With this feature, an application can provide a short reference recording of between five and 60 seconds so the generated speech can match the reference voice, without additional model training or fine-tuning. This is not unrestricted voice cloning: access is limited, and developers must obtain appropriate authorization and consent for the audio they submit. Voice-prompting access requires limited-access approval and is subject to Microsoft’s safeguards.

Performance and pricing

Microsoft reports approximately 225 milliseconds of model-inference latency for generating 45 seconds of audio. This is the clearest published performance characteristic in the supplied information and explains why the model is aimed at real-time interaction. It should not be treated as a guaranteed end-to-end response time or service-level commitment.

The Microsoft AI model page lists MAI-Voice-2-Flash at $15 per one million characters. This is a character-based price rather than a price per minute of audio. Actual Azure Speech charges, regional availability, service configuration, and preview terms may affect the cost of a deployed application. Teams should confirm current billing details in Azure before estimating production expenditure.

The cost and speed trade-off is straightforward: MAI-Voice-2-Flash is attractive when fast turn-taking has more value than the highest possible long-form fidelity. If an application generates large volumes of prerecorded narration, the related MAI-Voice-2 model may be worth evaluating even if its design target differs.

Integration and supported interface

MAI-Voice-2-Flash is integrated through Azure Speech and is also supported in Voice Live for real-time voice interactions. Developers need an Azure Speech resource in a supported region. The supplied documentation identifies SSML-based control, streaming, and audio generation as relevant capabilities.

Its supported modality profile is narrow and clear:

  • Input: Text.
  • Output: Generated audio speech.
  • Streaming: Supported.
  • Image, video, and audio input: Not documented as supported for this model.
  • Tool or function calling: Not documented as supported by the model.
  • Structured JSON output: Not documented as a model capability.

Voice Live can be used to build broader real-time voice experiences, but that does not make MAI-Voice-2-Flash a general reasoning or agent model. Application logic, conversation management, tools, and any language-model reasoning would need to be provided by the surrounding system.

Limits and considerations

MAI-Voice-2-Flash is a public-preview service. Microsoft does not provide a service-level agreement for the preview, and the model is not recommended for production workloads that require guaranteed service commitments. Availability and behavior can also depend on region and Azure service configuration.

No context-window size or maximum output-token limit is specified for this speech model. Those language-model measurements are not directly applicable to its documented text-to-speech role. The available information also does not identify an independent reasoning capability, coding capability, or model knowledge cutoff.

Voice and locale coverage is not completely uniform. Not every voice or locale supports every expressive style. The instant voice-prompting feature is gated, and applications must use consented reference material. Developers should also validate pronunciation, pacing, language switching, and emotional delivery with the exact voices and text patterns used in their product.

The model is optimized for responsiveness rather than maximum long-form fidelity. That makes it a less obvious choice for audiobooks, podcasts, extended narration, or other content where consistent speaker identity and the highest available audio quality are more important than quick interaction.

Best use cases

  • Real-time voice assistants: Fast spoken replies can make turn-taking feel more natural.
  • Conversational agents: Voice-based support or information systems can use expressive speech without waiting for long generation jobs.
  • Call-center automation: The model can provide spoken responses in customer-service workflows, subject to the surrounding system’s reliability and compliance requirements.
  • Interactive voice response: IVR applications can produce more natural prompts than fixed recordings when dynamic text is needed.
  • Multilingual applications: Support for 15 languages and 18 locales can help products serve users across multiple markets, provided the required voice and style are available.
  • Accessibility interfaces: Assistive applications can turn dynamic text into expressive speech.
  • Interactive narration and announcements: The low-latency design is useful when content is generated or changed at runtime.

When to choose MAI-Voice-2-Flash

Choose MAI-Voice-2-Flash when response speed is a central product requirement and the application needs expressive speech rather than a fixed collection of recordings. It is especially suitable when users are waiting for the next conversational turn, such as in a voice assistant, support call, or interactive IVR flow.

Choose a different option when the workload is primarily long-form narration and fidelity matters more than latency. Within Microsoft’s documented MAI-Voice family, MAI-Voice-2 is the relevant alternative for higher-fidelity extended generation. A different speech service may also be more appropriate if it offers the particular language, voice, compliance arrangement, availability guarantee, or pricing structure required by the deployment.

Overall, MAI-Voice-2-Flash is best understood as a specialized, fast speech-generation component. Its value comes from combining expressive output, multilingual coverage, streaming, and low reported inference latency. Its preview status, gated voice prompting, variable voice coverage, and lack of a service-level commitment are important constraints for anyone considering it beyond experimentation.


Answers to Frequently Asked Questions

When should I choose MAI-Voice-2-Flash instead of MAI-Voice-2?
Choose MAI-Voice-2-Flash when low response time and expressive speech are more important than maximum long-form fidelity, such as for real-time assistants, support calls, and interactive IVR. Choose the related MAI-Voice-2 model when producing extended narration, podcasts, audiobooks, or other content where consistent speaker quality and long-form fidelity matter more than immediate response speed.
How much does MAI-Voice-2-Flash cost?
The Microsoft AI model page lists MAI-Voice-2-Flash at $15 per one million characters. This is a character-based price rather than a per-minute audio rate. Actual Azure Speech costs may vary based on region, configuration, billing details, and preview terms.
Which languages and voice features does MAI-Voice-2-Flash support?
MAI-Voice-2-Flash supports 15 languages and 18 locales, including English, Italian, French, German, Hindi, Spanish, Portuguese, Korean, Simplified Chinese, Turkish, Russian, Thai, Dutch, Romanian, and Hungarian. It supports Microsoft’s prebuilt voices, SSML controls for delivery, and gated instant voice prompting using a five- to 60-second reference recording.
What is MAI-Voice-2-Flash?
MAI-Voice-2-Flash is Microsoft’s text-to-speech model available through Azure Speech in public preview. It converts text into expressive generated speech and is designed for applications that need fast spoken responses, such as voice assistants, conversational agents, and interactive voice response systems.
How fast is MAI-Voice-2-Flash?
Microsoft reports approximately 225 milliseconds of model-inference latency when generating 45 seconds of audio. This is a model-inference measurement, not a guaranteed end-to-end response time; network transfer, request processing, playback, and other service operations can add latency.


Sources 6
Provider

About Microsoft Copilot