MiMo-V2.5

MiMo-V2.5-ASR

by Xiaomi HyperAI · Current and accessible through the Xiaomi MiMo API; open-source code and weights available

MiMo-V2.5-ASR is Xiaomi MiMo’s specialized speech-to-text model for Chinese, English, regional Chinese dialects, code-switched speech, lyrics, noisy and far-field recordings, and multi-speaker audio. It supports API streaming, automatic language detection, MP3 and WAV input, duration-based pricing, and open-source local deployment.

Text Reasoning Coding
MiMo-V2.5-ASR is a specialized speech-to-text model from Xiaomi MiMo for audio that is more difficult than a clean, single-speaker recording. It supports Mandarin Chinese, English, several Chinese dialects, automatic language detection, bilingual code-switching, lyrics, overlapping speakers, background music, and automatic punctuation. Users can access it through Xiaomi’s API or deploy the published model locally.
Outputs

What MiMo-V2.5-ASR can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

3/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MiMo-V2.5
Model type Other
Release date 2026-06-02
Status Current and accessible through the Xiaomi MiMo API; open-source code and weights available
Knowledge cutoff notes

No model-specific knowledge-cutoff date was identified in the reviewed official Xiaomi documentation or official repository.

Model notes

MiMo-V2.5-ASR is a specialized automatic speech recognition model rather than a general conversational LLM. The API accepts a single audio input content part and documents MP3 and WAV formats. Language options are auto, Chinese, and English. Xiaomi reports support for Wu, Cantonese, Minnan, Sichuanese, and other dialects, as well as noisy, far-field, overlapping-speaker, lyrics, and knowledge-intensive audio. The model is available through the OpenAI-compatible Xiaomi MiMo API and has publicly released source code and weights. Context length, maximum output tokens, fine-tuning, caching, batch processing, JSON mode, and a knowledge-cutoff date are not specified in the reviewed official documentation. Editorial scores are comparative estimates for this specialized ASR model, not vendor benchmarks.

Cost

Model pricing

Input $0.074 per hour of input audio for overseas API usage; ¥0.5 per hour on the domestic pricing page
Output No separate output charge documented; billing is based on input-audio duration
Model guide

MiMo-V2.5-ASR: Xiaomi’s Open Speech Recognition Model for Difficult Audio

MiMo-V2.5-ASR is Xiaomi MiMo’s end-to-end automatic speech recognition model for transcribing Chinese, English, regional Chinese dialects, code-switched speech, lyrics, noisy recordings, far-field audio, and multi-speaker conversations. It is available through the Xiaomi MiMo API and as an open-source model with downloadable weights and code.

What is MiMo-V2.5-ASR?

MiMo-V2.5-ASR is an end-to-end automatic speech recognition model provided by Xiaomi MiMo. Its primary job is to convert spoken audio into written text. Unlike a general-purpose conversational model, it is designed specifically for transcription and related speech-processing workflows.

The model was released on June 2, 2026, as part of Xiaomi’s MiMo-V2.5 speech technology lineup. Xiaomi makes it available through the MiMo API and has also published source code and model weights for local experimentation and secondary development.

For a beginner, the practical distinction is simple: this model receives an audio recording and returns a transcript. It is intended for meetings, interviews, voice input, difficult recordings, lyrics, and other situations where accurately recognizing speech matters more than generating a conversational answer.

Recognition strengths and supported speech

MiMo-V2.5-ASR supports Mandarin Chinese, English, Chinese-English code-switching, and several regional Chinese varieties. Xiaomi specifically identifies Wu, Cantonese, Minnan, and Sichuanese among its supported dialects. The API can either detect the language automatically or receive an explicit Chinese or English language setting.

The model’s main distinction is its focus on challenging audio rather than only clean, close-microphone speech. Xiaomi positions it for:

  • Noisy recordings and far-field audio
  • Overlapping speech from multiple speakers
  • Conversations with background music
  • Chinese-English bilingual or code-switched speech
  • Lyrics in Chinese and English
  • Technical terms, names, place names, and classical poetry

Native punctuation generation is also supported. This means the returned transcript can include punctuation without requiring a separate punctuation-restoration step, although the final output should still be reviewed for domain-specific names, speaker attribution, and formatting requirements.

Where it fits in Xiaomi MiMo’s lineup

MiMo-V2.5-ASR occupies a specialized speech-recognition position within Xiaomi’s MiMo catalog. It is not a general chat model, a reasoning assistant, a coding model, or a speech-synthesis system. Its output is text transcription, and its input is audio.

This specialization explains both its usefulness and its limitations. A transcription model can be a better fit than a general multimodal model when the task is to process large numbers of recordings consistently and at an audio-duration price. However, it is not the appropriate choice for asking follow-up questions, generating software, creating images, producing video, or returning generated speech.

API access and input format

The Xiaomi MiMo API exposes MiMo-V2.5-ASR through an OpenAI-compatible chat-completions interface. Requests provide one audio content part containing Base64-encoded audio data. The documented supported formats are MP3 and WAV.

Streaming responses are available through server-sent events. Streaming can be useful when an application needs to display transcript text progressively rather than waiting for the complete recording to finish processing. The documented default limit for the ASR endpoint is 100 requests per minute and 10,000 tokens per minute, although actual limits may depend on account configuration and service conditions.

The API supports automatic language detection and explicit Chinese or English settings. The reviewed documentation describes an audio transcription input rather than a general-purpose multimodal request containing arbitrary combinations of text, images, video, and audio.

Pricing

Xiaomi’s overseas pay-as-you-go pricing lists MiMo-V2.5-ASR at $0.074 per hour of input audio. The domestic pricing page lists the model at ¥0.5 per hour. These prices are based on the duration of the input audio, not on the number of generated transcript tokens.

No separate output charge is documented in the supplied research. Because the pricing is duration-based, the model may be easier to estimate for batch transcription than a service billed primarily by text-token usage. Actual availability, currency, account requirements, and applicable regional pricing should be checked in the relevant Xiaomi MiMo account and pricing documentation before deployment.

Open-source availability and local deployment

MiMo-V2.5-ASR is also available as an open-source release with publicly accessible code and downloadable weights. This gives technical users an alternative to API-only access and allows them to investigate local deployment or build custom applications around the model.

The published repository identifies a Linux environment, Python 3.12, CUDA 12.0 or newer, the MiMo audio tokenizer, and additional dependencies including FlashAttention. It also provides a Gradio demonstration and a Python interface for transcription.

Local deployment offers more control over data handling and infrastructure, but it requires suitable hardware, software configuration, model downloads, and maintenance. The supplied research does not specify a minimum GPU memory requirement, supported hardware list, throughput benchmark, or a guaranteed local inference speed, so those details should not be assumed.

Technical limits and unavailable specifications

Several specifications commonly published for general language models are not identified in the reviewed Xiaomi documentation. The context length and maximum output-token limit are unspecified. A knowledge-cutoff date is also not provided, although that concept is less central to a transcription model than to a factual question-answering model.

There is no documented fine-tuning service, batch API, caching feature, JSON mode, or general tool/function-calling capability for MiMo-V2.5-ASR. The research also does not establish a separate speaker-diarization output format. Although the model is described as supporting overlapping multi-speaker conversations, users should not assume that it will automatically return reliably labeled speaker turns unless Xiaomi’s current implementation documents that behavior.

The model produces text rather than image, video, audio, or music output. It should therefore be treated as a speech-to-text component, not as a complete voice-agent stack. Applications requiring spoken replies would need a separate text-to-speech system, while applications requiring semantic search or question answering would need additional processing after transcription.

Reasoning, coding, tools, speed, and cost

Reasoning and coding are not the model’s intended capabilities. It may recognize technical vocabulary or transcribe speech about programming, but that is different from writing code, solving a multi-step problem, or validating an answer. Similarly, the model does not provide documented web search, external tools, or function calling.

The editorial assessment supplied with the research rates its relative reasoning capability at 3 out of 10 and coding capability at 1 out of 10. These are comparative editorial scores, not Xiaomi benchmarks or vendor-published ratings. The same assessment rates speed at 7 out of 10 and cost at 8 out of 10, reflecting the model’s specialized purpose and duration-based price rather than a claim of universal performance superiority.

In practical terms, MiMo-V2.5-ASR trades breadth for focus. A general multimodal model may offer more reasoning, vision, tool use, or conversational behavior, but a dedicated ASR model can be the more direct and economical choice when the required result is a transcript from supported audio.

Best use cases

MiMo-V2.5-ASR is a strong candidate for workflows where speech recognition is the central task and the audio may contain language variation or acoustic difficulty. Suitable examples include:

  • Meeting, interview, and lecture transcription
  • Chinese-English bilingual conversations
  • Recognition of supported Chinese dialects
  • Far-field recordings from rooms or distant microphones
  • Noisy conversations and recordings with background music
  • Multi-speaker discussion transcription
  • Chinese or English lyrics transcription
  • Voice input for applications that process the transcript downstream

For production use, teams should still evaluate representative recordings. Accuracy can vary with microphone quality, accents, overlapping speech, background noise, rare names, specialist vocabulary, and the desired punctuation or formatting style.

When to choose MiMo-V2.5-ASR

Choose MiMo-V2.5-ASR when the input is supported audio, the desired output is text, and the recording may include Chinese, English, dialect speech, code-switching, music, noise, distance from the microphone, or more than one speaker. Its API access, streaming support, open-source release, and audio-hour pricing make it relevant both for hosted applications and for teams investigating local deployment.

Another type of model may be more appropriate when the application needs general conversation, image or video understanding, coding, web research, structured tool use, or generated audio. A separate model or processing pipeline may also be preferable when the workflow requires verified speaker labels, a known maximum context size, formal batch processing, fine-tuning, or a documented structured-output format, because those capabilities are not specified for MiMo-V2.5-ASR in the supplied documentation.

Bottom line

MiMo-V2.5-ASR is a focused Xiaomi MiMo speech-recognition model for turning difficult audio into text. Its most useful differentiators are support for Chinese and English, selected Chinese dialects, code-switching, lyrics, noisy and far-field recordings, overlapping speakers, automatic language detection, punctuation, API streaming, and open-source access. Its limitations are equally important: it is not a general-purpose assistant, its context and output limits are undocumented, and several production features such as fine-tuning, batch processing, JSON mode, and tool use are not specified.


Answers to Frequently Asked Questions

What are the main limitations of MiMo-V2.5-ASR?
MiMo-V2.5-ASR is a speech-to-text model rather than a general assistant, coding model, reasoning system, or speech-generation tool. Its context length and maximum output-token limit are unspecified, and documented support for fine-tuning, batch processing, JSON mode, tool calling, and reliable speaker diarization is not provided.
How much does MiMo-V2.5-ASR cost?
Xiaomi’s overseas pay-as-you-go pricing lists MiMo-V2.5-ASR at $0.074 per hour of input audio. The domestic pricing page lists it at ¥0.5 per hour. The supplied documentation describes duration-based input pricing and does not specify a separate output charge.
How can developers access MiMo-V2.5-ASR?
Developers can access MiMo-V2.5-ASR through the Xiaomi MiMo API or use its open-source code and model weights for local deployment. The API uses an OpenAI-compatible chat-completions interface, accepts Base64-encoded MP3 or WAV audio, and supports streaming responses through server-sent events.
What is MiMo-V2.5-ASR used for?
MiMo-V2.5-ASR is an end-to-end automatic speech recognition model from Xiaomi MiMo that converts audio into written text. It is designed for transcription workflows such as meetings, interviews, lectures, voice input, lyrics, and difficult recordings.
Which languages and audio conditions does MiMo-V2.5-ASR support?
The model supports Mandarin Chinese, English, Chinese-English code-switching, and regional Chinese varieties including Wu, Cantonese, Minnan, and Sichuanese. It is intended for noisy and far-field audio, overlapping speakers, background music, technical terms, names, place names, and Chinese or English lyrics.


Sources 8
Provider

About Xiaomi HyperAI