What is MAI-Transcribe-1.5?
MAI-Transcribe-1.5 is a Microsoft automatic speech-recognition model. Automatic speech recognition, or ASR, converts spoken audio into written text. In practical terms, the model can take a recording of a meeting, interview, call, lecture, or other supported audio and return a transcript.
The model belongs to Microsoft’s MAI-Transcribe family and is listed as a preview model in Microsoft Foundry and Azure Speech. The supplied provider documentation identifies June 2, 2026 as its release date. It is not positioned as a general-purpose assistant: it does not generate images or video, synthesize speech, answer questions from a built-in knowledge cutoff, or serve as a normal text-generation model.
Microsoft’s documentation describes support for 43 languages, automatic language identification, noisy real-world audio, and keyword or entity biasing. Biasing lets an application provide terms that the recognizer should pay special attention to, which can be useful for company names, product names, technical vocabulary, medical terms, or other words that are easily misrecognized.
Where it fits in Microsoft’s catalog
MAI-Transcribe-1.5 sits in Microsoft’s speech and audio model catalog, not in the general-purpose conversational model category. Access is documented through Microsoft Foundry and Azure Speech. This positioning matters because the model is intended to be integrated into an application or audio-processing workflow rather than used as a standalone chat destination.
Microsoft describes MAI-Transcribe-1.5 as the second generation of the MAI-Transcribe family. The available comparison material also places it alongside a successor, MAI-Transcribe-2, but the relevant specifications here concern MAI-Transcribe-1.5. The model should therefore be evaluated as a focused transcription service: its value comes from audio recognition quality, language coverage, processing cost, and workflow compatibility rather than broad reasoning or content-generation features.
Key specifications
| Specification | MAI-Transcribe-1.5 |
|---|---|
| Provider | Microsoft |
| Model family | MAI-Transcribe |
| Status | Preview |
| Primary function | Speech-to-text transcription |
| Supported languages | 43, according to the supplied Microsoft materials |
| Audio input | Yes |
| Text output | Yes |
| Image, video, or audio output | No |
| Automatic language identification | Yes |
| Keyword or entity biasing | Up to 200 terms |
| Native real-time streaming | Not available in the current comparison |
| Speaker diarization | Not available in the current comparison |
| Word-level timestamps | Not supported according to the supplied notes |
| Documented context window | Not provided |
| Documented maximum output tokens | Not provided |
The missing context and token limits should not be interpreted as unlimited capacity. The supplied research does not provide a maximum audio duration, maximum transcript size, or maximum output-token specification for this model. Applications should verify the limits imposed by the particular Azure Speech or Microsoft Foundry interface they use.
Accuracy and processing speed
Microsoft reports an average FLEURS word error rate of 3.7% across its top 25 languages. The supplied research also records a 2.4% word error rate on Artificial Analysis. Word error rate, or WER, measures the number of word substitutions, deletions, and insertions relative to a reference transcript; lower is better. These are reported benchmark claims, not a guarantee of performance on every recording.
Real-world results can vary with microphone quality, background noise, accents, overlapping speech, specialist vocabulary, language switching, and recording conditions. Keyword or entity biasing can help with important terminology, but it does not eliminate the need to review transcripts in high-stakes settings.
Microsoft’s current comparison reports approximately 20 seconds of inference for one hour of audio in the relevant version comparison. That figure suggests a strong batch-processing speed advantage for workflows that do not require live captions. It should not be treated as a universal latency guarantee: processing time can depend on the endpoint, audio format, workload, service capacity, and application design.
Pricing and cost trade-offs
The supplied pricing information lists MAI-Transcribe-1.5 at $0.36 per hour of audio. No separate output charge is provided in the research. Because the model is billed by audio duration rather than by text tokens, its cost structure is easier to estimate for recorded-content workloads: an organization can approximate transcription expense from the number of audio hours processed.
The hourly pricing is especially relevant for large backlogs of meetings, interviews, calls, lectures, or media files. However, the listed price should be confirmed in the Microsoft Foundry or Azure pricing interface before deployment. Preview services can have changing availability, regional differences, quotas, or associated platform charges that are not represented by the model’s headline audio rate.
Compared with a larger general-purpose model that first transcribes audio and then performs additional reasoning, a dedicated ASR model may offer a more direct and economical path when the required output is simply a transcript. On the other hand, MAI-Transcribe-1.5 does not replace a language model for summarization, question answering, extraction, translation, or reasoning over the transcript; those tasks may require a separate processing stage.
Supported inputs and outputs
The model accepts audio and returns text. The supplied Microsoft Foundry material describes batch transcription access and supported audio formats, but the research does not provide a complete format matrix or a single universal duration limit. Developers should consult the selected Azure Speech or Foundry endpoint for the exact request requirements.
Automatic language identification is useful when the language is not known in advance or when an application processes recordings from multiple markets. The 43-language coverage is a major part of the model’s positioning, although the quality and feature availability may differ by language.
MAI-Transcribe-1.5 does not provide direct image, video, audio, or music output. Its output is text. It also is not documented as a tool-calling model, reasoning model, coding model, or structured-output model. The database’s editorial assessment rates its reasoning and coding usefulness at 1 out of 10, but these are editorial scores rather than Microsoft-published benchmark results. In practical terms, the model should not be selected to write software, plan complex tasks, invoke external functions, or generate JSON as its primary job.
Important limitations
- No native streaming: The current comparison lists native real-time streaming as unavailable. This makes the model better suited to recorded or batch audio than to applications requiring immediate partial transcripts.
- No speaker diarization: The supplied research says speaker diarization is unavailable. A transcript may therefore need a separate method or service if the application must identify who said each utterance.
- No word-level timestamps: The model is not documented as providing word-level timing. This is a limitation for precise subtitle alignment, searchable media players, and applications that need to synchronize every word with the audio.
- Not a general-purpose model: It does not generate text in the usual language-model sense, create images, produce video, synthesize speech, or answer general questions by itself.
- Preview status: Preview availability and behavior can change. Production teams should validate service-level expectations, regional availability, quotas, and API compatibility before committing to a long-lived integration.
- Unspecified capacity limits: No context length, maximum output-token count, or universal maximum audio duration is supplied. These limits should be checked in the specific service documentation and deployment path.
Best use cases
MAI-Transcribe-1.5 is a good fit when the central requirement is converting multilingual audio into text at batch-processing cost and speed. Suitable examples include:
- Transcribing recorded meetings, interviews, lectures, and research sessions.
- Creating first-pass captions or accessibility transcripts.
- Processing call-center recordings for later search, review, or analysis.
- Recognizing terminology from a known vocabulary using up to 200 biased keywords or entities.
- Building an audio-ingestion stage for a voice agent or content workflow.
- Converting large archives of noisy real-world recordings into searchable text.
For a production workflow, a common pattern would be to transcribe the audio with MAI-Transcribe-1.5 and then pass the resulting text to a separate model or rules-based system for summarization, classification, redaction, sentiment analysis, or structured extraction. That second stage is not a capability of the transcription model itself.
When to choose MAI-Transcribe-1.5
Choose MAI-Transcribe-1.5 when multilingual batch transcription, automatic language identification, noisy-audio handling, and predictable per-audio-hour pricing matter more than conversational reasoning. Its reported processing speed and $0.36-per-hour price make it particularly attractive for recorded audio at scale, subject to confirmation of current Azure pricing and service limits.
Consider another option when the application needs live partial results, speaker labels, word-level timestamps, or an all-in-one audio-to-insight system. A different speech service may be more appropriate if those features are mandatory. A general-purpose language model may be preferable after transcription when the main task is reasoning over the content, writing code, summarizing discussions, answering questions, or returning structured records. If the project requires speech synthesis, image understanding, video generation, or audio generation, MAI-Transcribe-1.5 is also the wrong component because its role ends at speech recognition.
Overall assessment
MAI-Transcribe-1.5 is a focused Microsoft speech-recognition model rather than a broad AI assistant. Its strongest documented characteristics are 43-language support, automatic language identification, keyword and entity biasing, reported low word error rates, fast batch processing, and a stated price of $0.36 per audio hour. Those features make it a practical candidate for multilingual transcription pipelines and large recorded-audio workloads.
Its limitations are equally important: it is preview software, lacks documented context and output limits, does not currently provide native streaming, and is not documented to provide diarization or word-level timestamps. The best evaluation question is therefore not whether it can replace a general AI model, but whether its transcription quality, language coverage, processing model, and cost match the specific audio workflow. For that narrow purpose, it can serve as an efficient first stage; additional tools will be needed for speaker attribution, real-time interaction, and higher-level analysis.

