MiMo-V2.6

MiMo-V2.6-Distill-Qwen-9B

by Xiaomi HyperAI · Current open-weight SFT checkpoint

MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo’s open-weight 9B model for coding, cybersecurity, visual coding, tool use and agentic research. Fine-tuned from Qwen3.5-9B on 77.4 billion tokens, it supports text and image input, produces text, uses a configured 262,144-token context length and is available under the MIT license. It can be deployed with tools such as Transformers, vLLM or SGLang, but users must provide their own infrastructure because no separate official hosted API price is specified.

Text Reasoning Coding
MiMo-V2.6-Distill-Qwen-9B is a compact open-weight model from Xiaomi MiMo aimed at developers who want an agentic model they can download and run themselves. It combines text generation with image understanding and is especially oriented toward coding, terminal-style tool use, cybersecurity tasks, visual coding and experimentation with agentic reinforcement learning. The published checkpoint is available under the MIT license, but it has no separate official hosted API price; users must provide their own compatible hardware or pay for third-party infrastructure.
Outputs

What MiMo-V2.6-Distill-Qwen-9B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

7/10 Reasoning
8/10 Coding
7/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family MiMo-V2.6
Model type Multimodal
Context window 262K tokens
Release date 2026-09-22
Status Current open-weight SFT checkpoint
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was published in the reviewed Xiaomi MiMo or Hugging Face documentation.

Model notes

MiMo-V2.6-Distill-Qwen-9B is a 9B model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on 77.4B tokens of MiMo-generated data. The released checkpoint is intended as a starting point for agentic reinforcement-learning research. It is available under the MIT license and is approximately 18.8 GB in the published bfloat16 safetensors format. The model card demonstrates image-text-to-text inference and local serving with Transformers, vLLM and SGLang. The configured maximum position length is 262,144 tokens. Benchmark results are vendor-reported and include internal evaluation sets. The checkpoint is distinct from the much larger hosted or open-sourced MiMo-V2.6-Pro and MiMo-V2.6-Flash models. No official knowledge cutoff, hosted API price, maximum generation limit, fine-tuning service, batch API, prompt-caching feature or separate JSON-mode capability was identified.

Cost

Model pricing

Input No official hosted API price; open-weight download
Output No official hosted API price; open-weight download
Model guide

MiMo-V2.6-Distill-Qwen-9B: Xiaomi’s Open Agentic Model for Local Coding

MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo’s open-weight 9-billion-parameter model for coding agents, tool-oriented workflows, cybersecurity experimentation, visual coding and general agent tasks. Fine-tuned from Qwen3.5-9B on 77.4 billion tokens of MiMo-generated data, it accepts text and images, produces text, supports a configured 262,144-token context length, and is released under the MIT license for local deployment and reinforcement-learning research.

What is MiMo-V2.6-Distill-Qwen-9B?

MiMo-V2.6-Distill-Qwen-9B is a 9-billion-parameter open-weight language model developed by Xiaomi MiMo. Xiaomi describes it as a supervised fine-tuning checkpoint built from Qwen3.5-9B. In practical terms, supervised fine-tuning means the base model was trained further on examples selected and prepared to improve particular behaviors. For this checkpoint, those behaviors focus on software engineering, cybersecurity, general agent tasks and visual tasks.

The model is intended to be downloaded and deployed through compatible inference software rather than used as a Xiaomi-hosted consumer chatbot. Its release is also positioned as a starting point for agentic reinforcement-learning research. Xiaomi provides associated reinforcement-learning environments, training code and lightweight agent harnesses, although the checkpoint itself remains the central item for users who want a relatively compact local model.

It sits below Xiaomi’s much larger MiMo-V2.6-Pro and MiMo-V2.6-Flash models in deployment scale. That positioning matters: the 9B checkpoint is more realistic for local experimentation than a large hosted model, but it should not automatically be expected to match the capabilities of Xiaomi’s larger variants on every task.

Training focus and intended purpose

The model card reports a weighted supervised fine-tuning mixture containing 77.4 billion total tokens, including 27.2 billion loss-bearing tokens. The training data is divided across four broad areas:

  • Software engineering: 23.2 billion total tokens
  • Cybersecurity: 11.0 billion total tokens
  • General agent tasks: 22.0 billion total tokens
  • Visual tasks: 21.2 billion total tokens

This distribution explains why the model is a better fit for coding agents and tool-oriented workflows than for a general-purpose chat experience alone. It can generate plans, code and tool-call text, interpret visual information and work through multi-step tasks. Tool use here means interaction through text-based agent protocols or an external harness; the model does not directly control a terminal, application or device without software connecting those tools to it.

Capabilities and supported modalities

MiMo-V2.6-Distill-Qwen-9B accepts text and images and produces text. Its multimodal support is therefore primarily image understanding: an application can provide an image alongside a prompt and ask the model to describe, analyze or reason about the visual content. The available research does not identify native image, audio, video, speech, music or embedding generation.

The model supports reasoning-oriented generation through the MiMo v2.6 chat template. When used with a suitable runtime, thinking can be enabled explicitly and reasoning content can be returned separately from the final answer. This makes it useful for tasks where the model needs to plan before producing a result, such as debugging a codebase, navigating a terminal workflow or analyzing a visual programming problem. The supplied research does not establish a separate official reasoning score or a guaranteed reasoning quality level.

Its main practical capabilities include:

  • Code generation, explanation and debugging
  • Agent planning and text-based tool interaction
  • Terminal and software-engineering workflows
  • Cybersecurity experimentation and analysis
  • Visual coding and image-assisted reasoning
  • General-purpose text generation and problem solving

These capabilities describe the model’s training emphasis and documented use cases, not a guarantee that every application will achieve the same results. Tool execution, file access and environment control depend on the surrounding agent framework.

Context length and deployment requirements

The published configuration specifies a maximum position length of 262,144 tokens. This is a configured context limit, meaning the model can theoretically process a very large amount of text within one request when the serving stack and available memory support it. It should not be confused with a guaranteed maximum output length: no authoritative model-specific maximum output-token limit was identified in the supplied documentation.

The released weights are provided in bfloat16 safetensors format and occupy approximately 18.8 GB on Hugging Face. That storage figure does not include runtime overhead, the operating system, the inference framework or the memory needed for the context and intermediate computations. A full-precision local deployment therefore requires substantially more practical capacity than the raw file size alone suggests. Quantization may reduce the hardware burden, but the supplied research does not specify a particular quantization format or performance level.

Official examples cover Transformers, vLLM, SGLang, Docker Model Runner, Google Colab and Kaggle. Xiaomi recommends a recent SGLang build with Qwen3.5 support and the MiMo reasoning parser for text generation. Users should check framework compatibility before deployment because support for the model’s multimodal architecture, chat template and reasoning output can vary between runtimes.

What the reported evaluations show

Xiaomi’s model card reports improvements over the Qwen3.5-9B base model on several agentic and coding evaluations. The following figures are vendor-reported; some use Xiaomi’s internal evaluation sets, so they are useful indicators of the model’s training focus rather than independent guarantees of production performance.

BenchmarkQwen3.5-9BMiMo-V2.6-Distill-Qwen-9B
SWE Verified60.061.1
SWE Pro32.044.6
MiMo Code mini19.551.6
MiMo Cyber mini5.731.3
Terminal Bench 2.127.037.1
Toolathlon-Verified25.935.2
MiMo Visual Coding mini61.764.0

The largest reported differences appear on the MiMo Code mini and MiMo Cyber mini evaluations, which is consistent with the checkpoint’s specialized training mixture. However, benchmark results can depend on prompting, tools, scaffolding and evaluation design. A developer should test the model with representative repositories, commands, images and security scenarios before relying on it in a production system.

Pricing, license and availability

MiMo-V2.6-Distill-Qwen-9B is an open-weight download rather than a separately priced hosted API model. Xiaomi publishes the checkpoint under the MIT license. There is consequently no official per-token input price, output price or subscription tier for this specific model in the supplied research.

“Free” in this context refers to the model license and download, not necessarily to deployment. Users may incur costs for GPUs, storage, electricity, cloud instances, inference hosting, bandwidth or a third-party service that makes the model available through an API. Xiaomi’s broader MiMo platform has separate developer services and pricing, but those should not be treated as the price of this open checkpoint.

Strengths and limitations

Strengths

  • Open deployment: The MIT license and downloadable weights allow local use, integration and research without depending on Xiaomi’s hosted endpoint.
  • Focused agent training: Coding, terminal work, cybersecurity and tool-oriented tasks are central to the training mixture.
  • Multimodal input: Image understanding adds value for visual coding, screenshots and other image-assisted workflows.
  • Large configured context: The 262,144-token position length can support long code or multi-step agent sessions when hardware and runtime support it.
  • Research suitability: The checkpoint is explicitly released as a starting point for agentic reinforcement-learning experiments.

Limitations

  • Local hardware burden: The approximately 18.8 GB bfloat16 weight files are only part of the memory requirement.
  • No managed API included: Users seeking predictable hosted inference, official per-token billing or provider-managed scaling must use another service or arrange hosting separately.
  • Text output only: It does not natively generate images, audio, video, speech or embeddings.
  • Unspecified limits: No authoritative maximum output-token limit, fine-tuning service, batch API, prompt-caching feature or separate JSON-mode capability was identified.
  • Evaluation caveats: The published benchmark results are vendor-reported and partly based on internal evaluation sets.
  • Runtime dependence: Reasoning output, image handling and tool workflows depend on the selected serving framework and external harness.

When to choose MiMo-V2.6-Distill-Qwen-9B

Choose this model when you want an openly licensed checkpoint for local coding agents, terminal automation, cybersecurity experimentation, visual coding or reinforcement-learning research. It is particularly attractive when control over deployment, data handling and inference software matters more than access to a managed API. Its 9B scale also makes it a more practical research target than Xiaomi’s much larger MiMo-V2.6 variants, although actual hardware requirements remain significant.

A hosted model may be more appropriate when you need simple integration, provider-managed scaling, predictable latency or official usage-based billing. A larger model may be preferable for difficult tasks where maximum reasoning quality is more important than local control and infrastructure cost. Conversely, a smaller or more heavily quantized model may be better when speed, memory usage or low-cost edge deployment is the priority.

MiMo-V2.6-Distill-Qwen-9B is not the best choice if the primary requirement is native media generation, speech processing, embeddings or a ready-made consumer assistant. It is also not a drop-in replacement for an agent platform: developers still need to supply the tools, permissions, execution environment, monitoring and safety controls around the model.

Bottom line

MiMo-V2.6-Distill-Qwen-9B is a focused open-weight model for developers who want coding and agent capabilities in a locally deployable 9B checkpoint. Its strongest differentiators are the training emphasis on software engineering and agent tasks, image understanding, a 262,144-token configured context length, MIT licensing and Xiaomi’s reported gains on several coding and cybersecurity evaluations. The trade-off is operational: users must handle hardware, serving, tool integration and evaluation themselves, and the model offers no separate official hosted API price or guaranteed maximum output limit.


Answers to Frequently Asked Questions

Is MiMo-V2.6-Distill-Qwen-9B free to use, and what license does it have?
MiMo-V2.6-Distill-Qwen-9B is available as an open-weight download under the MIT license. There is no separate official per-token price or subscription tier for this checkpoint, but users may still pay for GPUs, storage, electricity, cloud hosting, bandwidth or third-party inference services.
What hardware and deployment requirements does MiMo-V2.6-Distill-Qwen-9B have?
The bfloat16 safetensors weights occupy approximately 18.8 GB, excluding runtime overhead, context memory and intermediate computations. Official deployment examples include Transformers, vLLM, SGLang, Docker Model Runner, Google Colab and Kaggle. Quantization may reduce hardware requirements, but the supplied documentation does not specify a particular quantization format or performance level.
Does MiMo-V2.6-Distill-Qwen-9B support images and long context?
Yes. It accepts text and images and produces text, making it suitable for image understanding and visual coding tasks. Its published configuration specifies a maximum position length of 262,144 tokens, although actual usable context depends on the runtime and available hardware.
What is MiMo-V2.6-Distill-Qwen-9B?
MiMo-V2.6-Distill-Qwen-9B is a 9-billion-parameter open-weight language model developed by Xiaomi MiMo and fine-tuned from Qwen3.5-9B. It is designed for software engineering, cybersecurity, agent tasks, visual coding and image-assisted reasoning.
What can MiMo-V2.6-Distill-Qwen-9B be used for?
The model can support code generation, explanation and debugging, terminal workflows, text-based tool interaction, agent planning, cybersecurity experimentation and visual coding. Tool execution, file access and environment control require an external harness or agent framework.


Sources 4
Provider

About Xiaomi HyperAI