What is MiMo-V2.5-ASR?
MiMo-V2.5-ASR is an end-to-end automatic speech recognition model provided by Xiaomi MiMo. Its primary job is to convert spoken audio into written text. Unlike a general-purpose conversational model, it is designed specifically for transcription and related speech-processing workflows.
The model was released on June 2, 2026, as part of Xiaomi’s MiMo-V2.5 speech technology lineup. Xiaomi makes it available through the MiMo API and has also published source code and model weights for local experimentation and secondary development.
For a beginner, the practical distinction is simple: this model receives an audio recording and returns a transcript. It is intended for meetings, interviews, voice input, difficult recordings, lyrics, and other situations where accurately recognizing speech matters more than generating a conversational answer.
Recognition strengths and supported speech
MiMo-V2.5-ASR supports Mandarin Chinese, English, Chinese-English code-switching, and several regional Chinese varieties. Xiaomi specifically identifies Wu, Cantonese, Minnan, and Sichuanese among its supported dialects. The API can either detect the language automatically or receive an explicit Chinese or English language setting.
The model’s main distinction is its focus on challenging audio rather than only clean, close-microphone speech. Xiaomi positions it for:
- Noisy recordings and far-field audio
- Overlapping speech from multiple speakers
- Conversations with background music
- Chinese-English bilingual or code-switched speech
- Lyrics in Chinese and English
- Technical terms, names, place names, and classical poetry
Native punctuation generation is also supported. This means the returned transcript can include punctuation without requiring a separate punctuation-restoration step, although the final output should still be reviewed for domain-specific names, speaker attribution, and formatting requirements.
Where it fits in Xiaomi MiMo’s lineup
MiMo-V2.5-ASR occupies a specialized speech-recognition position within Xiaomi’s MiMo catalog. It is not a general chat model, a reasoning assistant, a coding model, or a speech-synthesis system. Its output is text transcription, and its input is audio.
This specialization explains both its usefulness and its limitations. A transcription model can be a better fit than a general multimodal model when the task is to process large numbers of recordings consistently and at an audio-duration price. However, it is not the appropriate choice for asking follow-up questions, generating software, creating images, producing video, or returning generated speech.
API access and input format
The Xiaomi MiMo API exposes MiMo-V2.5-ASR through an OpenAI-compatible chat-completions interface. Requests provide one audio content part containing Base64-encoded audio data. The documented supported formats are MP3 and WAV.
Streaming responses are available through server-sent events. Streaming can be useful when an application needs to display transcript text progressively rather than waiting for the complete recording to finish processing. The documented default limit for the ASR endpoint is 100 requests per minute and 10,000 tokens per minute, although actual limits may depend on account configuration and service conditions.
The API supports automatic language detection and explicit Chinese or English settings. The reviewed documentation describes an audio transcription input rather than a general-purpose multimodal request containing arbitrary combinations of text, images, video, and audio.
Pricing
Xiaomi’s overseas pay-as-you-go pricing lists MiMo-V2.5-ASR at $0.074 per hour of input audio. The domestic pricing page lists the model at ¥0.5 per hour. These prices are based on the duration of the input audio, not on the number of generated transcript tokens.
No separate output charge is documented in the supplied research. Because the pricing is duration-based, the model may be easier to estimate for batch transcription than a service billed primarily by text-token usage. Actual availability, currency, account requirements, and applicable regional pricing should be checked in the relevant Xiaomi MiMo account and pricing documentation before deployment.
Open-source availability and local deployment
MiMo-V2.5-ASR is also available as an open-source release with publicly accessible code and downloadable weights. This gives technical users an alternative to API-only access and allows them to investigate local deployment or build custom applications around the model.
The published repository identifies a Linux environment, Python 3.12, CUDA 12.0 or newer, the MiMo audio tokenizer, and additional dependencies including FlashAttention. It also provides a Gradio demonstration and a Python interface for transcription.
Local deployment offers more control over data handling and infrastructure, but it requires suitable hardware, software configuration, model downloads, and maintenance. The supplied research does not specify a minimum GPU memory requirement, supported hardware list, throughput benchmark, or a guaranteed local inference speed, so those details should not be assumed.
Technical limits and unavailable specifications
Several specifications commonly published for general language models are not identified in the reviewed Xiaomi documentation. The context length and maximum output-token limit are unspecified. A knowledge-cutoff date is also not provided, although that concept is less central to a transcription model than to a factual question-answering model.
There is no documented fine-tuning service, batch API, caching feature, JSON mode, or general tool/function-calling capability for MiMo-V2.5-ASR. The research also does not establish a separate speaker-diarization output format. Although the model is described as supporting overlapping multi-speaker conversations, users should not assume that it will automatically return reliably labeled speaker turns unless Xiaomi’s current implementation documents that behavior.
The model produces text rather than image, video, audio, or music output. It should therefore be treated as a speech-to-text component, not as a complete voice-agent stack. Applications requiring spoken replies would need a separate text-to-speech system, while applications requiring semantic search or question answering would need additional processing after transcription.
Reasoning, coding, tools, speed, and cost
Reasoning and coding are not the model’s intended capabilities. It may recognize technical vocabulary or transcribe speech about programming, but that is different from writing code, solving a multi-step problem, or validating an answer. Similarly, the model does not provide documented web search, external tools, or function calling.
The editorial assessment supplied with the research rates its relative reasoning capability at 3 out of 10 and coding capability at 1 out of 10. These are comparative editorial scores, not Xiaomi benchmarks or vendor-published ratings. The same assessment rates speed at 7 out of 10 and cost at 8 out of 10, reflecting the model’s specialized purpose and duration-based price rather than a claim of universal performance superiority.
In practical terms, MiMo-V2.5-ASR trades breadth for focus. A general multimodal model may offer more reasoning, vision, tool use, or conversational behavior, but a dedicated ASR model can be the more direct and economical choice when the required result is a transcript from supported audio.
Best use cases
MiMo-V2.5-ASR is a strong candidate for workflows where speech recognition is the central task and the audio may contain language variation or acoustic difficulty. Suitable examples include:
- Meeting, interview, and lecture transcription
- Chinese-English bilingual conversations
- Recognition of supported Chinese dialects
- Far-field recordings from rooms or distant microphones
- Noisy conversations and recordings with background music
- Multi-speaker discussion transcription
- Chinese or English lyrics transcription
- Voice input for applications that process the transcript downstream
For production use, teams should still evaluate representative recordings. Accuracy can vary with microphone quality, accents, overlapping speech, background noise, rare names, specialist vocabulary, and the desired punctuation or formatting style.
When to choose MiMo-V2.5-ASR
Choose MiMo-V2.5-ASR when the input is supported audio, the desired output is text, and the recording may include Chinese, English, dialect speech, code-switching, music, noise, distance from the microphone, or more than one speaker. Its API access, streaming support, open-source release, and audio-hour pricing make it relevant both for hosted applications and for teams investigating local deployment.
Another type of model may be more appropriate when the application needs general conversation, image or video understanding, coding, web research, structured tool use, or generated audio. A separate model or processing pipeline may also be preferable when the workflow requires verified speaker labels, a known maximum context size, formal batch processing, fine-tuning, or a documented structured-output format, because those capabilities are not specified for MiMo-V2.5-ASR in the supplied documentation.
Bottom line
MiMo-V2.5-ASR is a focused Xiaomi MiMo speech-recognition model for turning difficult audio into text. Its most useful differentiators are support for Chinese and English, selected Chinese dialects, code-switching, lyrics, noisy and far-field recordings, overlapping speakers, automatic language detection, punctuation, API streaming, and open-source access. Its limitations are equally important: it is not a general-purpose assistant, its context and output limits are undocumented, and several production features such as fine-tuning, batch processing, JSON mode, and tool use are not specified.

