MiMo V2.5

MiMo-V2.5-Pro

by Xiaomi HyperAI · Deprecated; API model name scheduled to expire on 2026-10-21 at 10:00 Beijing Time

Xiaomi MiMo-V2.5-Pro is a text-only reasoning model for repository-scale coding, long documents, and multi-step agent workflows. It offers a 1-million-token context window, 128,000-token maximum output, tool calls, web search, streaming, structured output, caching, and an MIT-licensed open-source release. Its hosted API identifier is scheduled to expire on October 21, 2026.

Text Reasoning Coding
MiMo-V2.5-Pro is Xiaomi’s flagship model in the MiMo V2.5 generation. It is built for tasks that require more than a short conversational answer: repository-scale coding, long documents, multi-step planning, tool-driven research, and workflows that may continue through hundreds of intermediate actions. The model accepts and produces text, with a 1-million-token context window and support for up to 128,000 output tokens. Xiaomi offers it through the MiMo API and has released the MiMo-V2.5 series as open source under the MIT license. However, Xiaomi has announced that the API identifier mimo-v2.5-pro will be deprecated on October 21, 2026, with no direct replacement specified.
Outputs

What MiMo-V2.5-Pro can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Fine-tuning Structured output Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiMo V2.5
Model type Reasoning
Context window 1M tokens
Maximum output 128K tokens
Knowledge cutoff December 2024
Release date 2026-04-23
Status Deprecated; API model name scheduled to expire on 2026-10-21 at 10:00 Beijing Time
Deprecation date 2026-10-21
Shutdown date 2026-10-21
Knowledge cutoff notes

Xiaomi's official API example for mimo-v2.5-pro states that the model's knowledge cutoff date is December 2024. This is presented in the documented system-message example rather than in a separate model specification field.

Model notes

Canonical API identifier is mimo-v2.5-pro. Xiaomi describes the model as a trillion-parameter architecture with approximately 42B active parameters, 1M context, deep thinking, tool calls, web search, structured output, streaming, and context caching. The model is text-only. Xiaomi open-sourced the MiMo-V2.5 series under the MIT license on 2026-06-29 and states that commercial inference deployment and secondary training are permitted. Fine-tuning is therefore marked positive for the open-source deployment path, but a first-party hosted fine-tuning endpoint was not verified. Xiaomi's deprecation notice states that mimo-v2.5-pro will expire on 2026-10-21 at 10:00 Beijing Time, with no system replacement model. The official API example states a December 2024 knowledge cutoff.

Cost

Model pricing

Input ¥0.025 per million tokens cache hit; ¥3 per million tokens cache miss; equivalent USD prices $0.0036 and $0.435 per million tokens
Output ¥6 per million tokens; equivalent USD price $0.87 per million tokens
Model guide

MiMo-V2.5-Pro: Xiaomi’s Long-Context Model for Coding and Autonomous Agents

MiMo-V2.5-Pro is Xiaomi’s text-only, trillion-parameter model for complex software engineering, long-horizon agent workflows, tool use, and large documents. Its defining specifications are a 1-million-token context window and a maximum output of 128,000 tokens. It supports web search, tool calls, streaming, structured output, and context caching, and the MiMo-V2.5 series is available under the MIT license for open-source deployment. Xiaomi has announced that the hosted API model name will expire on October 21, 2026, so its deprecation status is an important consideration for new projects.

What is MiMo-V2.5-Pro?

MiMo-V2.5-Pro is Xiaomi’s agent-oriented foundation model for demanding text-based workloads. In practical terms, it is intended for applications where the model must understand a large amount of information, reason through a sequence of decisions, write or modify code, and use external tools before producing a final result.

Its primary audience is developers building complex agent systems and software-engineering workflows. Xiaomi’s examples include compiler development, full-stack application work, and electronic-design automation. These are tasks where a model may need to inspect many files, maintain context across a long process, call tools repeatedly, and revise its work rather than answer in a single exchange.

MiMo-V2.5-Pro is provided by Xiaomi through the MiMo model platform. It occupies the flagship, high-capability position in the MiMo V2.5 series rather than serving as a lightweight chat model. That positioning helps explain its unusually large context and output limits, but also means that it is not automatically the best choice for every short or latency-sensitive request.

Core specifications and context capacity

SpecificationVerified detail
ProviderXiaomi
Model identifiermimo-v2.5-pro
Model familyMiMo V2.5
Input and outputText in, text out
Context window1,000,000 tokens
Maximum output128,000 tokens
Architecture descriptionApproximately 1 trillion total parameters and 42 billion active parameters
Knowledge cutoffDecember 2024, according to Xiaomi’s documented API example
LicenseMIT for the open-source MiMo-V2.5 release

A token is a small unit of text used by a language model. The 1-million-token context window is therefore not the same as a one-million-word limit, but it is still large enough for very extensive source repositories, long technical documentation, multiple research records, or a lengthy multi-step interaction. The context window includes the material supplied to the model and the conversation or tool information retained for the request.

The 128,000-token output ceiling is also unusually high. It can be useful when the model must generate a substantial codebase, a detailed technical artifact, or a long intermediate result. In normal use, requesting the maximum output would not necessarily be efficient: longer generations cost more and can take longer to review.

Reasoning, coding, and agent workflows

MiMo-V2.5-Pro is classified as a reasoning model in the supplied model data. Its intended advantage is sustained problem solving: breaking a broad objective into steps, examining intermediate results, and continuing after tool responses. Xiaomi describes the model as supporting deep thinking and sustained execution across nearly 1,000 tool calls.

Tool calls allow an application to expose defined functions or services to the model. For example, a coding agent could provide functions for listing files, reading source code, running tests, or applying a patch. The model can then decide when to invoke those functions and use their results in later reasoning. The tool-call count described by Xiaomi is a provider claim about the model’s demonstrated or supported workflow capability, not a guarantee that every application will achieve the same result.

The model is particularly suited to software engineering because it combines coding capability with a very large context and tool support. A repository-scale workflow might involve reading architectural documentation, tracing dependencies, editing several files, running tests, interpreting compiler errors, and making follow-up changes. MiMo-V2.5-Pro is designed for this kind of extended loop rather than only for producing isolated code snippets.

The available data rates its reasoning and coding capabilities highly, but those ratings are editorial evaluations rather than Xiaomi-published benchmark scores. They should be treated as guidance about intended positioning, not as independently verified performance measurements.

Supported features and modalities

MiMo-V2.5-Pro is text-only. It accepts text input and produces text output; native image, audio, video, speech, music, and other non-text output are not documented for this model. An application can still use external tools to process files or connect to other services, but that does not make the model itself a native multimodal generator.

Documented platform features include:

  • Tool calls: integration with application-defined functions and external services.
  • Web search: access to a search capability through the platform, where enabled by the application.
  • Streaming: delivery of generated output incrementally instead of waiting for the complete response.
  • Structured output: generation constrained or organized into a specified structure for application processing.
  • Context caching: reuse of context in supported workflows to reduce repeated processing costs.

Structured output is documented for the model, but a separate, distinct JSON-mode capability has not been verified in the supplied research. Developers should therefore distinguish between asking for a structured response and relying on a specifically documented JSON-only mode.

Pricing and cost trade-offs

Xiaomi lists pay-as-you-go pricing for MiMo-V2.5-Pro by token usage. The documented Chinese prices are ¥0.025 per million tokens for cache hits, ¥3 per million tokens for cache misses, and ¥6 per million output tokens. The same pricing documentation gives equivalent U.S. dollar figures of $0.0036 per million cache-hit tokens, $0.435 per million cache-miss tokens, and $0.87 per million output tokens.

Cache-hit pricing is substantially lower than cache-miss pricing. This makes context caching relevant for applications that repeatedly use the same long instructions, repository material, or other persistent context. Developers should still account for the cost of newly supplied context and generated output, particularly when an agent performs many tool calls or produces long intermediate responses.

The model’s cost and speed trade-off are important. Its large context, reasoning orientation, and long-horizon workflows make it more suitable for complex tasks than for every short request. The supplied editorial assessment gives it a speed score of 7 out of 10 and a cost score of 9 out of 10; these are editorial scores, not provider specifications. In practice, a smaller or faster model may be more appropriate for simple classification, short drafting, routine extraction, or high-volume interactions where the full capability of MiMo-V2.5-Pro is unnecessary.

Open-source release and deployment

Xiaomi announced that the MiMo-V2.5 series was open-sourced under the MIT license on June 29, 2026. The release permits commercial inference deployment and secondary training according to Xiaomi’s documentation. Xiaomi also documented adaptations involving vLLM, SGLang, AWS Trainium2, AMD ROCm, and other accelerator ecosystems.

This open-source availability gives organizations an alternative to using only the hosted API. Self-managed deployment can offer greater control over infrastructure, data handling, and serving configuration, but it also requires suitable hardware, model-serving expertise, and responsibility for operational reliability. The supplied research supports marking fine-tuning as available for the open-source deployment path, but it does not verify a first-party hosted fine-tuning endpoint.

Limitations and deprecation status

The most significant current limitation is lifecycle risk. Xiaomi has announced that the hosted API model name mimo-v2.5-pro will expire on October 21, 2026, at 10:00 Beijing Time. Requests using that identifier are expected to fail after the deadline. The notice does not specify a system replacement, so applications should not assume that requests will automatically route to a newer MiMo model.

This does not necessarily make the open-source release unusable, but it does make the hosted API a poor fit for a new production system that cannot accommodate migration. Teams evaluating the model should separate two decisions: whether MiMo-V2.5-Pro is technically suitable, and whether its announced API lifetime is acceptable.

The model also has no native image, audio, video, speech, or music generation. Applications needing those capabilities should use a model designed for the relevant modality or connect MiMo-V2.5-Pro to separate services. Its knowledge cutoff is documented as December 2024, so current information should come through web search or other tools rather than being assumed from the model’s internal knowledge.

When to choose MiMo-V2.5-Pro

MiMo-V2.5-Pro is a strong candidate when the task has several of the following characteristics:

  • The application must process very large documents or repositories.
  • The workflow requires planning and multiple intermediate steps.
  • The model needs to call tools repeatedly and respond to their results.
  • The main workload involves complex software engineering or coding agents.
  • Long outputs or substantial generated artifacts are useful.
  • Context caching can reduce the cost of repeatedly supplied material.
  • An MIT-licensed, self-deployable model is preferable to a hosted-only service.

Another type of model may be more appropriate for short answers, low-latency interaction, simple extraction, routine classification, or multimodal generation. A smaller model can also be the better operational choice when the task does not benefit from a million-token context or extended reasoning. Finally, applications requiring a stable hosted endpoint beyond October 21, 2026 should choose an option with a confirmed longer support period or complete a migration plan before adopting this identifier.

Overall assessment

MiMo-V2.5-Pro is defined less by ordinary chat use than by its combination of long context, coding ability, tool use, and extended agent execution. Its 1-million-token context and 128,000-token output limit give it a practical role in repository-scale engineering and long-document workflows, while web search, streaming, structured output, and caching support production integrations.

Its text-only design limits multimodal use, and its announced API deprecation is a material concern for new deployments. For teams that can use the open-source release or plan around the lifecycle deadline, it is a technically focused option for complex agents. For simpler, faster, multimodal, or long-lived hosted workloads, another model type may be a better fit.


Answers to Frequently Asked Questions

When will the hosted MiMo-V2.5-Pro API be deprecated?
Xiaomi has announced that the hosted API identifier `mimo-v2.5-pro` will expire on October 21, 2026, at 10:00 Beijing Time. Applications using the hosted endpoint should plan a migration unless they can use the open-source release for self-managed deployment.
How much does MiMo-V2.5-Pro cost?
Xiaomi’s documented pay-as-you-go pricing is ¥0.025 per million tokens for cache hits, ¥3 per million tokens for cache misses, and ¥6 per million output tokens. The equivalent listed U.S. prices are $0.0036, $0.435, and $0.87 per million tokens, respectively.
What features does MiMo-V2.5-Pro support?
The model supports text input and text output, tool calls, web search, streaming, structured output, and context caching. It is text-only and does not natively generate or understand images, audio, video, speech, or music.
How large is MiMo-V2.5-Pro’s context window and maximum output?
MiMo-V2.5-Pro supports a 1,000,000-token context window and a maximum output of 128,000 tokens. This capacity is intended for large code repositories, extensive documentation, multi-step tool interactions, and substantial generated artifacts.
What is MiMo-V2.5-Pro designed for?
MiMo-V2.5-Pro is Xiaomi’s agent-oriented foundation model for complex text-based workloads, including repository-scale software engineering, compiler development, full-stack applications, electronic-design automation, long-document analysis, and workflows that require repeated tool calls.


Sources 6
Provider

About Xiaomi HyperAI