MiMo-V2-Flash

MiMo-V2-Flash-Base

by Xiaomi HyperAI · Current open-weight base model

Xiaomi MiMo-V2-Flash-Base is a 309B-total, 15B-active open-weight mixture-of-experts foundation model with a 256K context window. It is available for local deployment through Hugging Face and supported inference engines, but no official hosted API pricing or maximum output limit was identified for the base checkpoint.

Text Reasoning Coding
MiMo-V2-Flash-Base is the base checkpoint in Xiaomi’s MiMo-V2-Flash family. Its mixture-of-experts architecture, hybrid attention design and multi-token-prediction components target efficient long-context language modeling. The weights are available under the MIT license, but Xiaomi has not documented official hosted pricing, a maximum output-token limit or a first-party API endpoint for this exact checkpoint.
Outputs

What MiMo-V2-Flash-Base can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MiMo-V2-Flash
Model type General Purpose
Context window 256K tokens
Knowledge cutoff December 2024
Release date 2025-12-16
Status Current open-weight base model
Knowledge cutoff notes

The MiMo-V2-Flash repository includes December 2024 as the knowledge cutoff in its recommended English and Chinese system prompts. This is not presented as a separate model-card metadata field.

Model notes

MiMo-V2-Flash-Base is the base checkpoint associated with Xiaomi's MiMo-V2-Flash series. Official materials describe a 309B-total-parameter MoE model with 15B active parameters, hybrid sliding-window/global attention, 27T training tokens, and a 256K context window. The weights are published on Hugging Face under the MIT license and can be run with Transformers, vLLM, or SGLang using trusted remote code. The model is text-generation-only in the published model card; no native image, audio, or video input/output is documented. The December 2024 knowledge cutoff appears in Xiaomi's recommended system prompt and should be treated as a stated model cutoff rather than a separately published technical specification. No official hosted pricing, maximum output-token limit, fine-tuning service, batch API, JSON mode, or first-party web-search integration was verified for this exact base checkpoint. Editorial scores are comparative estimates, not vendor benchmarks.

Model guide

MiMo-V2-Flash-Base: Xiaomi’s Open-Weight 309B MoE Foundation Model

MiMo-V2-Flash-Base is Xiaomi’s open-weight 309-billion-parameter mixture-of-experts language model, with 15 billion active parameters and a documented 256K-token context window. It is designed primarily for research, custom post-training and self-hosted long-context inference rather than as a turnkey hosted chat API.

MiMo-V2-Flash-Base is Xiaomi’s open-weight foundation model for users who want to experiment with, adapt or self-host a large language model. It is not positioned in the supplied documentation as a ready-made consumer chatbot or a conventionally priced hosted API model. Instead, it is the base checkpoint of Xiaomi’s MiMo-V2-Flash series, intended for research, custom post-training and deployment through compatible inference software.

The model combines a very large total parameter count with a much smaller number of active parameters per token. That design aims to provide the representational capacity of a large model while reducing the computation required for each individual prediction. Its documented 256K-token context window also makes it relevant to workloads involving long documents, large codebases or extended research material, provided the operator has sufficient hardware and a suitable inference setup.

What is MiMo-V2-Flash-Base?

MiMo-V2-Flash-Base is a text-generation model released by Xiaomi under the XiaomiMiMo organization. It is described as a 309-billion-parameter mixture-of-experts, or MoE, model with 15 billion active parameters. In an MoE model, the full parameter set is divided into specialized expert components, and the routing system selects only part of that network for a particular token. “309B” therefore describes the model’s total parameter capacity, while “15B active” describes the approximate amount used for an individual token’s computation.

This distinction matters when evaluating deployment. The active-parameter figure can help explain the model’s computational efficiency relative to a dense model of the same total size, but it does not mean that the model requires only the memory of a 15-billion-parameter checkpoint. The complete published weights still represent a very large model, so self-hosting requires substantial infrastructure and careful quantization or parallelism decisions.

The “Base” designation is also important. A base checkpoint is a foundation model intended to be adapted or used as a starting point, rather than a model specifically optimized as a polished instruction-following assistant. The supplied research does not establish that this checkpoint offers the conversational behavior, tool calling or structured-output guarantees that a hosted instruction model might provide.

Verified specifications and lineup position

MiMo-V2-Flash-Base was listed with a release date of December 16, 2025, and is identified as a current open-weight base model. It belongs to Xiaomi’s MiMo-V2-Flash family. Xiaomi’s official repository and model card describe a model trained on 27 trillion tokens and using hybrid sliding-window and global attention. These architectural choices are intended to balance local attention efficiency with the ability to retain access to information across a long sequence.

SpecificationDocumented detail
ProviderXiaomi
Model familyMiMo-V2-Flash
Model roleOpen-weight base foundation model
Total parameters309 billion
Active parameters15 billion
Context window256,000 tokens
Training data volume27 trillion tokens, as stated in the official materials
LicenseMIT
Knowledge cutoffDecember 2024, stated in the recommended system prompts
Primary outputText

The December 2024 knowledge cutoff should be treated carefully. It appears in Xiaomi’s recommended English and Chinese system prompts, rather than as a separately documented model-card metadata field. It is therefore a stated model prompt detail, not necessarily a complete technical description of every item in the training data.

Long-context design and practical meaning

The documented 256K-token context window is one of the model’s clearest practical differentiators. A context window is the amount of text the model can consider in a single request, including the prompt and the generated continuation. A 256K-token limit can accommodate substantially more material than the context windows of many smaller models, although the actual usable length depends on the inference engine, memory configuration and the format of the request.

For example, the model may be suitable for examining a large technical repository, comparing several lengthy documents, or maintaining continuity across a long research session. A long context does not guarantee perfect recall of every detail. Retrieval quality, prompt organization and the model’s ability to identify relevant information remain important, especially when the input approaches the upper limit.

Xiaomi’s materials describe hybrid sliding-window and global attention. In simple terms, sliding-window attention lets the model concentrate efficiently on nearby tokens, while global attention provides routes for information to remain available across much longer sequences. The model also includes multi-token-prediction components. The supplied research identifies these components as part of the design but does not provide a verified benchmark showing how much they improve throughput or generation quality in a particular deployment.

Modalities, reasoning and coding capabilities

The published model card identifies MiMo-V2-Flash-Base as a text-generation-only model. No native image, audio or video input or output is documented. It should therefore be evaluated as a language model that accepts text and produces text, not as a multimodal assistant. Images, audio files and video cannot be assumed to work without an external preprocessing system and an additional model.

The model can be relevant to reasoning and coding experiments because it is a general-purpose language foundation model with a large context window. However, the supplied sources do not establish a dedicated reasoning mode, a provider-published reasoning benchmark or special reasoning controls for this base checkpoint. The editorial reasoning score of 8 out of 10 is a comparative estimate, not a Xiaomi-published measurement.

The same distinction applies to coding. The model may be used for code generation, explanation, transformation and repository-level experimentation, but the documentation does not present a dedicated coding model or verified coding benchmark for this exact checkpoint. The editorial coding score of 8 out of 10 reflects an assessment of likely usefulness, not an official capability rating.

MiMo-V2-Flash-Base is not documented as having native function calling, tool use, web search or code execution. It may be possible for an operator to build an external application that interprets generated text and connects it to tools, but that would be application-layer orchestration rather than a verified built-in feature of the model. JSON mode and guaranteed structured output are likewise not documented.

Deployment and inference options

The weights are published on Hugging Face and can be run with Transformers, vLLM or SGLang using trusted remote code, according to the supplied model notes. This gives technical users several routes for local or private deployment, but it does not make the model lightweight. The full 309-billion-parameter checkpoint remains a substantial deployment project.

MiMo-V2-Flash-Base is therefore better suited to organizations, researchers and experienced developers who can manage model files, GPU memory, distributed inference and performance tuning. Inference speed will depend heavily on hardware, quantization, parallelism, batch size and the length of the prompt. Xiaomi’s materials do not supply a universal tokens-per-second figure, so no single speed claim should be applied to every deployment.

The active-parameter architecture may improve the cost and speed trade-off compared with a similarly sized dense model, but the model should not automatically be considered inexpensive to operate. Weight storage, memory overhead, communication between devices and the long-context workload can all affect total cost. The editorial speed score of 9 out of 10 and cost score of 8 out of 10 are comparative estimates rather than provider benchmarks.

Pricing and API availability

No official hosted pricing was verified for MiMo-V2-Flash-Base. The supplied research also does not identify a first-party hosted endpoint, maximum output-token limit, fine-tuning service or batch API for this exact base checkpoint. As a result, there is no reliable per-token input or output price to report.

This is a significant difference from a managed API model. With a hosted service, the provider normally handles hardware, scaling and software updates in exchange for usage charges. With MiMo-V2-Flash-Base, the open-weight license gives users more control over deployment and customization, but the operator must provide the infrastructure and absorb its associated costs.

The absence of an official output limit should not be interpreted as unlimited generation. Every deployment will still be constrained by the configured context window, available memory, inference-server settings and any stopping conditions chosen by the operator.

Main strengths and limitations

  • Long-context capacity: The documented 256K-token context window supports unusually large prompts and document collections.
  • Open deployment: The weights are available under the MIT license, giving researchers and developers broad freedom to inspect, adapt and host the model subject to the license terms.
  • MoE efficiency: The 309B-total and 15B-active design aims to combine high capacity with lower per-token computation than a dense model of comparable total size.
  • Research flexibility: As a base checkpoint, it can serve as a foundation for custom post-training and experimentation rather than locking users into one hosted interaction format.
  • Infrastructure demands: Open weights do not remove the hardware and operational requirements associated with a 309-billion-parameter model.
  • Limited product guarantees: There is no verified official API price, maximum output setting, uptime commitment or managed scaling option for this checkpoint.
  • Text only: Native image, audio and video capabilities are not documented.
  • Base-model behavior: Users seeking a polished assistant may need additional instruction tuning, prompting or application-level controls.
  • No verified built-in tools: Function calling, web search, code execution, JSON mode and structured-output guarantees are not documented.

When to choose MiMo-V2-Flash-Base

Choose MiMo-V2-Flash-Base when you need an open-weight foundation model for long-context research, custom post-training or self-hosted inference. It is particularly relevant when control over the weights, deployment environment and adaptation process matters more than having a simple hosted API. Potential projects include long-document analysis, codebase experimentation, model fine-tuning research and evaluation of MoE inference strategies.

It is also a reasonable candidate when a 256K-token context is materially useful and your team can support the associated infrastructure. The model’s combination of high total capacity and fewer active parameters may be attractive for experimentation with large-model quality and inference efficiency, although the final trade-off must be measured on your own hardware and workload.

When another option may be more appropriate

A managed hosted model may be a better choice if you need predictable per-request pricing, automatic scaling, simple integration, service-level guarantees or a documented API contract. An instruction-tuned conversational model may also be more suitable for a customer-facing assistant that needs reliable dialogue behavior without additional post-training.

A multimodal model is the better option for applications that must directly understand images, audio or video. Likewise, a model or platform with verified function calling and structured-output support is preferable when the application must reliably invoke business tools or return machine-readable data. MiMo-V2-Flash-Base can potentially be placed inside such a system, but the supplied research does not show that these capabilities are built into this checkpoint.

Overall assessment

MiMo-V2-Flash-Base is best understood as a large, open and adaptable research foundation model rather than a finished AI service. Its most concrete advantages are the 256K-token context window, MIT-licensed weights, MoE architecture and support for established self-hosting frameworks. Its main costs are operational: the model is large, text-only and not accompanied by verified hosted pricing or product-level guarantees.

For users evaluating open-weight models, the decision should center on deployment control, long-context requirements and the ability to operate a very large checkpoint. For users who mainly want a ready-to-use assistant, multimodal interaction or a clearly priced API, another type of model will likely be more practical.


Answers to Frequently Asked Questions

Does MiMo-V2-Flash-Base support multimodal input, function calling or an official API?
The documented checkpoint is text-only, with no verified native image, audio or video support. Function calling, web search, code execution, JSON mode and guaranteed structured output are also not documented as built-in features. No official hosted pricing or first-party API endpoint was verified for this exact base model.
Can MiMo-V2-Flash-Base be run locally or self-hosted?
Yes. The weights are available on Hugging Face under the MIT license and can be used with Transformers, vLLM or SGLang with trusted remote code. However, the complete 309-billion-parameter checkpoint requires substantial infrastructure, and deployment may involve quantization, distributed inference and GPU-memory optimization.
What is MiMo-V2-Flash-Base?
MiMo-V2-Flash-Base is Xiaomi’s open-weight 309-billion-parameter mixture-of-experts foundation model. It has approximately 15 billion active parameters per token and is intended for research, custom post-training, experimentation and self-hosted deployment rather than use as a finished consumer chatbot.
How large is MiMo-V2-Flash-Base’s context window?
MiMo-V2-Flash-Base has a documented 256,000-token context window. This can support long documents, large codebases and extended research sessions, although practical usage depends on the inference engine, available memory and prompt organization.


Sources 4
Provider

About Xiaomi HyperAI