What is MAI-Transcribe-2?
MAI-Transcribe-2 is a speech-to-text model from Microsoft AI. Its job is narrowly defined: it listens to an audio recording and produces a written transcript, optionally enriched with language information, speaker labels, confidence data, offsets, durations, and word-level timestamps.
The model is available in public preview through Azure Speech and Microsoft Foundry. In the Azure Speech transcription API, it is selected through enhanced mode with enhancedMode.enabled set to true and enhancedMode.model set to MAI-Transcribe-2. It is therefore best understood as a specialized transcription service rather than a general-purpose conversational AI model.
MAI-Transcribe-2 is intended for recordings made in realistic conditions: accented speech, background noise, inconsistent microphone quality, overlapping speakers, and conversations that move between languages. It is not a text chatbot, image model, speech synthesizer, embedding model, or media-generation system.
What the model can do
- Transcribe 60 languages: Microsoft documents support for English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Turkish, Ukrainian, Vietnamese, Cantonese, and other European, Asian, Middle Eastern, and African languages.
- Identify languages automatically: Language identification can be used when a locale is not supplied explicitly.
- Separate speakers: Speaker diarization labels different voices in a recording so that a meeting transcript can distinguish participants.
- Return word-level timestamps: Timing for individual words supports captions, audio search, transcript navigation, editing, and alignment with video.
- Use keyword biasing: Recognition can be guided toward domain-specific phrases, names, abbreviations, and specialist terminology.
- Handle code switching: The model is designed for speech that changes between languages, including examples such as Hinglish and Spanglish.
- Offer clean or verbatim output: Verbatim transcription preserves fillers, false starts, and self-corrections. Clean transcription removes disfluencies to produce more readable text.
These features make the model more suitable for operational transcription than a basic speech-recognition endpoint that returns only plain text. For example, a media team can use timestamps to align captions, while a contact center can use speaker attribution and keyword hints to search conversations for product names or compliance terms.
Supported audio, limits, and output
The supplied model information lists WAV, MP3, and FLAC as supported input formats. The model card specifies a maximum input of 300 MB or two hours, making it appropriate for many meetings, interviews, lessons, calls, and podcast segments.
The native output is text. Optional metadata can include speaker identity, detected language, offsets, durations, confidence values, and timestamps. Word-level timing is particularly useful when the transcript must remain synchronized with an audio or video source.
There is no documented context-window or maximum-output-token specification because MAI-Transcribe-2 does not operate as a conventional text-generation language model. Its practical output limit is tied to the supported audio input and the resulting transcript returned by the speech service, not to a published token context window.
Microsoft documents an important preview limitation: diarization requests for recordings of approximately 15 minutes or longer may time out or fail, even when the same audio transcribes successfully without speaker separation. For long recordings, Microsoft recommends disabling diarization and combining word-level timestamps with a separate speaker-diarization process.
Accuracy and speed claims
Microsoft reports an average FLEURS word error rate of 5.2% across 60 languages in its launch announcement. The model page also reports a 3.4% average WER across its top 25 languages and a 2% Artificial Analysis WER result. Word error rate measures transcription mistakes, so lower values indicate fewer word-level errors, but results can vary with language, accent, noise, overlapping speech, microphone quality, and recording conditions.
Microsoft also reports that one hour of audio can be processed in approximately 10 seconds under the cited inference conditions. This is a provider-reported performance claim rather than a guarantee for every Azure region, workload, file, or request configuration.
The supplied catalog assigns MAI-Transcribe-2 a speed score of 9 and a cost score of 10. Those are editorial comparison scores, not Microsoft benchmarks or service guarantees. They reflect the model’s positioning as a fast, inexpensive transcription option, particularly at its introductory price.
Pricing and access
The launch price is $0.10 per hour of audio as a limited-time offer available through December 31, 2026, according to the supplied Microsoft research. The price is based on audio duration rather than text output, and no separate text-output charge is documented in the provided information.
Using the model requires an Azure subscription, a Microsoft Foundry Speech resource, and an available Azure region. Actual access can also depend on preview availability and the configuration of the selected Azure service. Because the model is in public preview, users should verify current regional availability and pricing before committing it to a production workflow.
Reasoning, coding, and tool capabilities
MAI-Transcribe-2 does not provide general-purpose reasoning or coding capabilities. It recognizes and structures spoken language; it does not independently analyze a business problem, write software, browse the web, call external tools, generate embeddings, or carry on a general chat.
Its useful “intelligence” is task-specific. It can identify speech content, distinguish speakers when diarization works, detect languages, apply recognition hints, and produce either a readable or highly faithful transcript. Any later summarization, question answering, translation, coding, or workflow automation would normally require a separate model or application layer.
Main strengths and limitations
Strengths
- Broad multilingual coverage: Sixty supported languages and code-switching support make it suitable for international meetings and multilingual media.
- Useful transcript structure: Speaker labels and word-level timestamps support captions, search, editing, compliance review, and analytics.
- Domain adaptation: Keyword biasing can improve recognition of names, product terms, acronyms, and specialized vocabulary.
- Flexible transcript style: Clean output favors readability, while verbatim output preserves details important for research, legal review, or conversation analysis.
- Low introductory cost: The announced $0.10-per-hour price is attractive for high-volume transcription if the offer and regional availability apply to the workload.
- Designed for imperfect recordings: Microsoft positions the model for noise, accents, variable microphones, and other conditions found outside studio-quality audio.
Limitations
- Public-preview status: The service is not covered by a service-level agreement, so production-critical users should account for availability and behavior changes.
- Long-recording diarization risk: Speaker separation may fail or time out at approximately 15 minutes or longer, even when transcription without diarization succeeds.
- Limited input formats: The supplied specifications list WAV, MP3, and FLAC rather than a broad set of video or container formats.
- Not a general AI assistant: It cannot replace a language model for summarization, reasoning, coding, web research, or conversational question answering.
- Accuracy is condition-dependent: Published WER figures are benchmark or provider-reported results, not guaranteed error rates for every language and recording.
- No audio generation: The model outputs text and does not synthesize speech, music, or other audio.
Best use cases
MAI-Transcribe-2 is a strong fit when the primary requirement is accurate, structured transcription rather than open-ended interaction. Suitable applications include meeting records, contact-center analysis, clinical documentation, legal and financial recordings, accessibility captions, video subtitles, media archiving, podcast indexing, e-learning, multilingual content operations, and voice-agent evaluation.
It is especially useful when a transcript needs to answer questions such as “which speaker said this?”, “where in the recording was this phrase used?”, or “which language was spoken at this point?” Keyword biasing can help organizations handle terminology that ordinary recognition systems may misinterpret.
When to choose MAI-Transcribe-2
Choose MAI-Transcribe-2 when you need multilingual audio transcription with timestamps, speaker attribution, language identification, and recognition hints, and when a low per-hour price and high processing speed matter. It is also a sensible choice for teams already using Azure Speech or Microsoft Foundry and willing to work with a public-preview service.
Consider another option when diarization must work reliably on long recordings, when a production workflow requires an SLA, or when the application needs summarization, reasoning, coding, web access, or speech generation in the same model. For long meetings, a practical alternative based on Microsoft’s documented guidance is to transcribe without diarization and apply a separate speaker-identification stage afterward.
Bottom line
MAI-Transcribe-2 is a focused Azure speech-recognition model rather than a general AI system. Its combination of 60-language coverage, code-switching support, word-level timestamps, speaker diarization, keyword biasing, and clean or verbatim modes gives it a useful role in transcription pipelines. The main trade-off is preview maturity: the price and reported speed are compelling, but long-recording diarization failures and the absence of an SLA make validation essential before using it for critical workloads.

