MiMo-V2.6

MiMo-V2.6-Pro UltraSpeed

by Xiaomi HyperAI · Current; high-speed inference variant of MiMo-V2.6-Pro

MiMo-V2.6-Pro UltraSpeed is Xiaomi MiMo’s latency-focused variant of the MiMo-V2.6-Pro reasoning model. It accepts text, images, video, and audio, produces text, supports a 1-million-token context window, tool calling, streaming, web search, structured output, and context caching, and is priced separately for interactive workloads that benefit from faster inference.

Text Reasoning Coding
MiMo-V2.6-Pro UltraSpeed is Xiaomi’s ultra-fast deployment mode for the MiMo-V2.6-Pro model family. It is aimed at interactive applications where response time matters, including coding assistants, research tools, computer-operation workflows, and multimodal analysis. The model retains the Pro family’s 1-million-token context window, reasoning features, tool calling, streaming, web search, structured output, and context caching, but does not generate images, video, audio, music, or speech.
Outputs

What MiMo-V2.6-Pro UltraSpeed can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
10/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family MiMo-V2.6
Model type Reasoning
Context window 1M tokens
Maximum output 131K tokens
Release date 2026-09-22
Status Current; high-speed inference variant of MiMo-V2.6-Pro
Knowledge cutoff notes

Xiaomi's sample system prompt for the related MiMo-V2.6-Pro endpoint states a December 2024 knowledge cutoff, but the official UltraSpeed documentation does not independently specify a cutoff for the exact UltraSpeed identifier.

Model notes

The canonical API identifier is mimo-v2.6-pro-ultraspeed. Xiaomi describes UltraSpeed as a mode of MiMo-V2.6-Pro that can deliver up to 20x inference speed for highly interactive use cases. The Pro family documentation lists text, image, video, and audio input, text output, a 1M-token context window, tool calling, streaming, web search, structured output, and context caching. Xiaomi lists separate mainland China pricing of ¥0.25 per million cached-input tokens, ¥30 per million uncached-input tokens, and ¥60 per million output tokens. UltraSpeed does not support the batch API, and its rate limits are listed as customized services. Editorial scores are comparative estimates, not provider-published ratings.

Cost

Model pricing

Input $0.036 per million cached-input tokens; $4.35 per million uncached-input tokens
Output $8.70 per million output tokens
Model guide

MiMo-V2.6-Pro UltraSpeed: Xiaomi’s Low-Latency Multimodal Reasoning Model

MiMo-V2.6-Pro UltraSpeed is Xiaomi MiMo’s high-speed inference variant of the MiMo-V2.6-Pro reasoning model. It accepts text, images, video, and audio as input while producing text responses, and is designed for latency-sensitive coding, research, tool use, office automation, and multimodal agent workflows.

What is MiMo-V2.6-Pro UltraSpeed?

MiMo-V2.6-Pro UltraSpeed is a high-speed inference variant of Xiaomi MiMo’s MiMo-V2.6-Pro reasoning model. In practical terms, it is intended to provide the same general class of long-context, multimodal understanding and agent-oriented capabilities as the Pro model while prioritizing lower response latency.

Xiaomi describes UltraSpeed as a mode that can deliver up to 20 times the inference speed of the standard model in highly interactive scenarios. That figure is a provider claim rather than an independent benchmark result, so actual performance will depend on the workload, endpoint, network conditions, request size, and service availability.

The model has the canonical API identifier mimo-v2.6-pro-ultraspeed. Xiaomi also lists it as a selectable model in MiMo Desktop, meaning it is not only an internal performance setting but a separately identifiable deployment option in the current MiMo catalog.

Where it fits in Xiaomi MiMo’s lineup

UltraSpeed belongs to the MiMo-V2.6 model family and is specifically related to MiMo-V2.6-Pro. Its distinction is deployment speed rather than a separate output modality or a fundamentally different product category. The standard Pro model is positioned for complex projects, coding, research, cybersecurity, design, office work, and computer-operation tasks; UltraSpeed targets similar work when users need faster interaction.

This positioning makes it different from a lightweight text-only model. It is still designed for difficult reasoning and multimodal understanding, but the trade-off is higher standard API pricing than the regular MiMo-V2.6-Pro endpoint. Applications should therefore use UltraSpeed when latency has meaningful practical value rather than selecting it automatically for every request.

Supported inputs, outputs, and modalities

MiMo-V2.6-Pro UltraSpeed supports multiple input formats:

  • Text
  • Images
  • Video
  • Audio

Its native output is text. Multimodal input does not mean that the model can create a new image, video, audio recording, music track, or speech file. For example, it can analyze an image, summarize a video, interpret an audio recording, or combine those inputs with written instructions, but the documented output capability is text generation.

This makes UltraSpeed suitable for applications such as reviewing visual documents, extracting information from media, answering questions about recordings, and combining long written context with files or other media. It is not the appropriate choice when the central requirement is native image, video, audio, music, or speech generation.

Reasoning, coding, and agent features

UltraSpeed is a reasoning model intended for tasks that require more than short, direct answers. Xiaomi’s Pro family documentation positions it for long-horizon work, coding, research, cybersecurity, design, office tasks, and computer-operation workflows. The available research also rates its reasoning and coding suitability highly on an editorial comparative scale, but those scores are assessments rather than Xiaomi-published benchmark results.

The model supports tool calling, which allows an application to provide functions that the model can request during a conversation. This can support workflows such as retrieving information, invoking business tools, operating software, or coordinating multi-step tasks. Tool calling does not mean the model independently has unrestricted access to external systems; the application must define and execute the available tools.

UltraSpeed also supports first-party web search, streaming responses, structured output, and context caching. Streaming lets an application display text as it is produced instead of waiting for the complete response. Structured output helps applications request responses in a defined format, while context caching can reduce repeated processing costs or latency when the same context is reused, subject to Xiaomi’s API behavior and pricing rules.

Context window and maximum output

The MiMo-V2.6-Pro family supports a context window of up to 1 million tokens. A context window is the amount of conversation, instructions, documents, and other input the model can consider within a request or ongoing interaction. A 1-million-token window is particularly relevant to large codebases, extensive research material, lengthy records, and multi-document workflows.

Xiaomi’s model documentation lists a maximum output of 128K tokens. Its compatible Anthropic Messages API documentation specifies a maximum max_tokens range of 131,072 for the UltraSpeed identifier. These figures are close but not identical because they come from different documentation contexts. Implementations should follow the limit enforced by the specific endpoint being used rather than assuming that every interface exposes exactly the same maximum.

A large context window does not guarantee that every large prompt will be equally useful or inexpensive. Applications still need to manage document selection, prompt structure, output length, and the cost of uncached input tokens.

Pricing and API availability

Xiaomi’s overseas API pricing for MiMo-V2.6-Pro UltraSpeed is:

Usage typePrice
Cached input$0.036 per million tokens
Uncached input$4.35 per million tokens
Output$8.70 per million tokens

For mainland China, Xiaomi lists pricing of ¥0.25 per million cached-input tokens, ¥30 per million uncached-input tokens, and ¥60 per million output tokens. The prices above are token-usage rates, not a monthly consumer subscription price.

UltraSpeed is available through Xiaomi MiMo’s API and compatible OpenAI-style and Anthropic-style interfaces. Xiaomi does not support the batch API for this model. The documented rate-limit arrangement is customized rather than expressed as a standard published requests-per-minute or tokens-per-minute quota. This may matter for teams planning large-scale or asynchronous workloads.

Speed, cost, and capability trade-offs

The main reason to choose UltraSpeed is responsiveness. Faster inference can make an assistant feel more usable when a person is waiting for an answer, reviewing code interactively, directing a tool-using agent, or carrying out a sequence of computer-operation steps.

The trade-off is price. The uncached input and output rates are substantially higher than the cached-input rate, and UltraSpeed is priced separately from the regular MiMo-V2.6-Pro model. A workload that repeatedly processes the same large instructions or documents may benefit from caching, while a workload dominated by new input and long generated responses may incur higher costs.

UltraSpeed may also be unnecessary for jobs where speed is not important. For offline analysis, lower-priority document processing, or large asynchronous workloads, an alternative model or a batch-capable endpoint may be more appropriate. Xiaomi’s documentation specifically states that UltraSpeed does not support batch processing.

Best use cases

  • Interactive coding: Fast code explanations, debugging discussions, planning, and iterative software-engineering assistance.
  • Tool-using agents: Workflows that repeatedly reason, call a function, inspect the result, and continue.
  • Multimodal analysis: Questions about images, video, audio, and text when the required answer is written.
  • Long-context research: Comparing or interrogating large collections of documents while retaining extensive working context.
  • Computer-operation assistance: Interactive desktop and office workflows where delays accumulate across multiple steps.
  • Real-time drafting and planning: Writing, summarization, research planning, and content-production tasks where users need quick iteration.

When to choose MiMo-V2.6-Pro UltraSpeed

Choose UltraSpeed when the application needs Xiaomi MiMo’s Pro-level multimodal reasoning and long context but the user experience depends on low latency. It is especially suitable when each response is part of an interactive loop, such as a coding session, research assistant, tool-calling agent, or desktop workflow.

Consider the regular MiMo-V2.6-Pro option when the task can tolerate more delay and the lower standard price is more important than responsiveness. Consider a batch-capable alternative for large asynchronous jobs, because UltraSpeed does not support Xiaomi’s batch API. Finally, select a model with native media generation when the application must create images, video, audio, music, or speech rather than analyze media and return text.

Limitations and practical considerations

UltraSpeed should be understood as a fast text-output reasoning model with multimodal input, not as an all-purpose media-generation system. Its 1-million-token context window is useful for large inputs, but actual request limits and maximum output behavior can vary by interface. The API documentation should be checked for the selected endpoint, especially when using the Anthropic-compatible interface.

Pricing is another important limitation. The model’s speed advantage may justify the additional cost for interactive applications, but it may not be economical for every high-volume workload. Customized rate limits may also require deployment planning or coordination with Xiaomi for applications that need predictable throughput.

Overall, MiMo-V2.6-Pro UltraSpeed is best represented as a distinct deployable model identifier within the MiMo-V2.6-Pro family: a latency-focused choice for demanding multimodal reasoning, coding, long-context work, and agent interaction, with text output and higher usage pricing in exchange for speed.


Answers to Frequently Asked Questions

When should developers choose MiMo-V2.6-Pro UltraSpeed?
Choose MiMo-V2.6-Pro UltraSpeed when low latency is important in interactive coding, multimodal analysis, research, tool-using agents, or computer-operation workflows. The regular MiMo-V2.6-Pro may be more economical when delays are acceptable, while a batch-capable model is better for large asynchronous workloads because UltraSpeed does not support Xiaomi’s batch API.
How much does MiMo-V2.6-Pro UltraSpeed cost?
Xiaomi’s overseas API pricing is $0.036 per million cached-input tokens, $4.35 per million uncached-input tokens, and $8.70 per million output tokens. Mainland China pricing is ¥0.25 per million cached-input tokens, ¥30 per million uncached-input tokens, and ¥60 per million output tokens.
What are the context window and output limits of MiMo-V2.6-Pro UltraSpeed?
The MiMo-V2.6-Pro family supports a context window of up to 1 million tokens. Xiaomi lists a maximum output of 128K tokens, while the compatible Anthropic Messages API documentation specifies a maximum max_tokens value of 131,072. The enforced limit may depend on the selected endpoint.
What is MiMo-V2.6-Pro UltraSpeed?
MiMo-V2.6-Pro UltraSpeed is Xiaomi MiMo’s low-latency inference variant of the MiMo-V2.6-Pro reasoning model. It is designed for multimodal understanding, long-context tasks, coding, research, and agent workflows, with a focus on faster responses. Its canonical API identifier is "mimo-v2.6-pro-ultraspeed".
What input and output modalities does MiMo-V2.6-Pro UltraSpeed support?
The model accepts text, images, video, and audio as inputs. Its native output is text, so it can analyze or summarize media but does not natively generate images, video, audio, music, or speech.


Sources 6
Provider

About Xiaomi HyperAI