MAI-Transcribe

MAI-Transcribe-2

by Microsoft Copilot · Public preview

MAI-Transcribe-2 is Microsoft’s public-preview speech-to-text model for WAV, MP3, and FLAC audio. It supports 60 languages, automatic language identification, code switching, speaker diarization, word-level timestamps, keyword biasing, and clean or verbatim transcription styles. The model is available through Azure Speech and Microsoft Foundry at a limited-time launch price of $0.10 per hour, but diarization may fail on recordings of approximately 15 minutes or longer.

Text
MAI-Transcribe-2 is Microsoft’s specialized model for turning real-world audio into searchable, time-aligned transcripts. Available through Azure Speech and Microsoft Foundry, it is designed for multilingual recordings, noisy environments, changing speakers, and conversations that switch between languages. Its strongest practical advantages are broad language coverage, speaker labels, word-level timing, and a low introductory price, while its public-preview status and documented long-recording diarization problems are important limitations.
Outputs

What MAI-Transcribe-2 can produce

Text
Inputs

What it can understand

Audio
Model profile

Performance characteristics

9/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family MAI-Transcribe
Model type Other
Release date 2026-09-03
Status Public preview
Knowledge cutoff notes

Microsoft does not publish a separate knowledge cutoff for this speech-to-text model. Its behavior is based on audio recognition rather than a documented text knowledge date.

Model notes

MAI-Transcribe-2 is a specialized speech-recognition model rather than a general-purpose language model. It supports 60 languages, automatic language identification, code switching, speaker diarization, word-level timestamps, keyword biasing, and verbatim or clean transcription styles. Supported input formats are WAV, MP3, and FLAC. The model card specifies a maximum input of 300 MB or two hours. Azure documentation identifies the service as public preview and notes that diarization may fail or time out for recordings of approximately 15 minutes or longer. The model is accessed through Azure Speech or Microsoft Foundry using enhanced mode. The speed and cost scores are editorial comparisons for transcription models, not vendor-provided ratings.

Cost

Model pricing

Input $0.10 per hour of audio; limited-time launch pricing through December 31, 2026
Output Included; no separate text-output charge documented
Model guide

MAI-Transcribe-2: Microsoft’s Fast Multilingual Speech-to-Text Model

Microsoft MAI-Transcribe-2 is a public-preview speech-recognition model that converts WAV, MP3, and FLAC audio into time-aligned text across 60 languages. It combines automatic language identification, speaker diarization, word-level timestamps, keyword biasing, code-switching support, and clean or verbatim transcription styles, with launch pricing of $0.10 per hour through December 31, 2026.

What is MAI-Transcribe-2?

MAI-Transcribe-2 is a speech-to-text model from Microsoft AI. Its job is narrowly defined: it listens to an audio recording and produces a written transcript, optionally enriched with language information, speaker labels, confidence data, offsets, durations, and word-level timestamps.

The model is available in public preview through Azure Speech and Microsoft Foundry. In the Azure Speech transcription API, it is selected through enhanced mode with enhancedMode.enabled set to true and enhancedMode.model set to MAI-Transcribe-2. It is therefore best understood as a specialized transcription service rather than a general-purpose conversational AI model.

MAI-Transcribe-2 is intended for recordings made in realistic conditions: accented speech, background noise, inconsistent microphone quality, overlapping speakers, and conversations that move between languages. It is not a text chatbot, image model, speech synthesizer, embedding model, or media-generation system.

What the model can do

  • Transcribe 60 languages: Microsoft documents support for English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Turkish, Ukrainian, Vietnamese, Cantonese, and other European, Asian, Middle Eastern, and African languages.
  • Identify languages automatically: Language identification can be used when a locale is not supplied explicitly.
  • Separate speakers: Speaker diarization labels different voices in a recording so that a meeting transcript can distinguish participants.
  • Return word-level timestamps: Timing for individual words supports captions, audio search, transcript navigation, editing, and alignment with video.
  • Use keyword biasing: Recognition can be guided toward domain-specific phrases, names, abbreviations, and specialist terminology.
  • Handle code switching: The model is designed for speech that changes between languages, including examples such as Hinglish and Spanglish.
  • Offer clean or verbatim output: Verbatim transcription preserves fillers, false starts, and self-corrections. Clean transcription removes disfluencies to produce more readable text.

These features make the model more suitable for operational transcription than a basic speech-recognition endpoint that returns only plain text. For example, a media team can use timestamps to align captions, while a contact center can use speaker attribution and keyword hints to search conversations for product names or compliance terms.

Supported audio, limits, and output

The supplied model information lists WAV, MP3, and FLAC as supported input formats. The model card specifies a maximum input of 300 MB or two hours, making it appropriate for many meetings, interviews, lessons, calls, and podcast segments.

The native output is text. Optional metadata can include speaker identity, detected language, offsets, durations, confidence values, and timestamps. Word-level timing is particularly useful when the transcript must remain synchronized with an audio or video source.

There is no documented context-window or maximum-output-token specification because MAI-Transcribe-2 does not operate as a conventional text-generation language model. Its practical output limit is tied to the supported audio input and the resulting transcript returned by the speech service, not to a published token context window.

Microsoft documents an important preview limitation: diarization requests for recordings of approximately 15 minutes or longer may time out or fail, even when the same audio transcribes successfully without speaker separation. For long recordings, Microsoft recommends disabling diarization and combining word-level timestamps with a separate speaker-diarization process.

Accuracy and speed claims

Microsoft reports an average FLEURS word error rate of 5.2% across 60 languages in its launch announcement. The model page also reports a 3.4% average WER across its top 25 languages and a 2% Artificial Analysis WER result. Word error rate measures transcription mistakes, so lower values indicate fewer word-level errors, but results can vary with language, accent, noise, overlapping speech, microphone quality, and recording conditions.

Microsoft also reports that one hour of audio can be processed in approximately 10 seconds under the cited inference conditions. This is a provider-reported performance claim rather than a guarantee for every Azure region, workload, file, or request configuration.

The supplied catalog assigns MAI-Transcribe-2 a speed score of 9 and a cost score of 10. Those are editorial comparison scores, not Microsoft benchmarks or service guarantees. They reflect the model’s positioning as a fast, inexpensive transcription option, particularly at its introductory price.

Pricing and access

The launch price is $0.10 per hour of audio as a limited-time offer available through December 31, 2026, according to the supplied Microsoft research. The price is based on audio duration rather than text output, and no separate text-output charge is documented in the provided information.

Using the model requires an Azure subscription, a Microsoft Foundry Speech resource, and an available Azure region. Actual access can also depend on preview availability and the configuration of the selected Azure service. Because the model is in public preview, users should verify current regional availability and pricing before committing it to a production workflow.

Reasoning, coding, and tool capabilities

MAI-Transcribe-2 does not provide general-purpose reasoning or coding capabilities. It recognizes and structures spoken language; it does not independently analyze a business problem, write software, browse the web, call external tools, generate embeddings, or carry on a general chat.

Its useful “intelligence” is task-specific. It can identify speech content, distinguish speakers when diarization works, detect languages, apply recognition hints, and produce either a readable or highly faithful transcript. Any later summarization, question answering, translation, coding, or workflow automation would normally require a separate model or application layer.

Main strengths and limitations

Strengths

  • Broad multilingual coverage: Sixty supported languages and code-switching support make it suitable for international meetings and multilingual media.
  • Useful transcript structure: Speaker labels and word-level timestamps support captions, search, editing, compliance review, and analytics.
  • Domain adaptation: Keyword biasing can improve recognition of names, product terms, acronyms, and specialized vocabulary.
  • Flexible transcript style: Clean output favors readability, while verbatim output preserves details important for research, legal review, or conversation analysis.
  • Low introductory cost: The announced $0.10-per-hour price is attractive for high-volume transcription if the offer and regional availability apply to the workload.
  • Designed for imperfect recordings: Microsoft positions the model for noise, accents, variable microphones, and other conditions found outside studio-quality audio.

Limitations

  • Public-preview status: The service is not covered by a service-level agreement, so production-critical users should account for availability and behavior changes.
  • Long-recording diarization risk: Speaker separation may fail or time out at approximately 15 minutes or longer, even when transcription without diarization succeeds.
  • Limited input formats: The supplied specifications list WAV, MP3, and FLAC rather than a broad set of video or container formats.
  • Not a general AI assistant: It cannot replace a language model for summarization, reasoning, coding, web research, or conversational question answering.
  • Accuracy is condition-dependent: Published WER figures are benchmark or provider-reported results, not guaranteed error rates for every language and recording.
  • No audio generation: The model outputs text and does not synthesize speech, music, or other audio.

Best use cases

MAI-Transcribe-2 is a strong fit when the primary requirement is accurate, structured transcription rather than open-ended interaction. Suitable applications include meeting records, contact-center analysis, clinical documentation, legal and financial recordings, accessibility captions, video subtitles, media archiving, podcast indexing, e-learning, multilingual content operations, and voice-agent evaluation.

It is especially useful when a transcript needs to answer questions such as “which speaker said this?”, “where in the recording was this phrase used?”, or “which language was spoken at this point?” Keyword biasing can help organizations handle terminology that ordinary recognition systems may misinterpret.

When to choose MAI-Transcribe-2

Choose MAI-Transcribe-2 when you need multilingual audio transcription with timestamps, speaker attribution, language identification, and recognition hints, and when a low per-hour price and high processing speed matter. It is also a sensible choice for teams already using Azure Speech or Microsoft Foundry and willing to work with a public-preview service.

Consider another option when diarization must work reliably on long recordings, when a production workflow requires an SLA, or when the application needs summarization, reasoning, coding, web access, or speech generation in the same model. For long meetings, a practical alternative based on Microsoft’s documented guidance is to transcribe without diarization and apply a separate speaker-identification stage afterward.

Bottom line

MAI-Transcribe-2 is a focused Azure speech-recognition model rather than a general AI system. Its combination of 60-language coverage, code-switching support, word-level timestamps, speaker diarization, keyword biasing, and clean or verbatim modes gives it a useful role in transcription pipelines. The main trade-off is preview maturity: the price and reported speed are compelling, but long-recording diarization failures and the absence of an SLA make validation essential before using it for critical workloads.


Answers to Frequently Asked Questions

How much does MAI-Transcribe-2 cost?
The announced introductory price is $0.10 per hour of audio through December 31, 2026. Access requires an Azure subscription and a Microsoft Foundry Speech resource, and actual availability and pricing may vary by Azure region and preview status.
What are the main limitations of MAI-Transcribe-2?
MAI-Transcribe-2 is a public-preview transcription service without a documented service-level agreement. Its speaker diarization may fail on longer recordings, accuracy varies with language and recording conditions, and it does not provide general-purpose reasoning, summarization, coding, web access, or speech generation.
Can MAI-Transcribe-2 identify speakers and handle multilingual conversations?
Yes. The model supports speaker diarization, automatic language identification, and code-switching between languages, including conversational patterns such as Hinglish and Spanglish. However, diarization requests for recordings of approximately 15 minutes or longer may time out or fail during the public preview.
What is MAI-Transcribe-2?
MAI-Transcribe-2 is Microsoft AI’s specialized speech-to-text model for converting audio into written transcripts. It can also provide language detection, speaker labels, confidence data, offsets, durations, and word-level timestamps.
Which languages and audio formats does MAI-Transcribe-2 support?
MAI-Transcribe-2 supports 60 languages, including English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Turkish, Ukrainian, Vietnamese, and Cantonese. The documented audio formats are WAV, MP3, and FLAC, with input limited to 300 MB or two hours.


Sources 6
Provider

About Microsoft Copilot