MiMo-VL

MiMo-VL-7B-SFT-2508

by Xiaomi HyperAI · Current open-weight checkpoint; available for download

MiMo-VL-7B-SFT-2508 is Xiaomi MiMo’s MIT-licensed, open-weight 7B vision-language checkpoint for image and video understanding, multimodal reasoning, and further supervised fine-tuning or reinforcement learning. It supports a 128K-token context and text output, but has no verified official hosted API price or native tool integrations.

Text Reasoning Coding
MiMo-VL-7B-SFT-2508 is an open-weight multimodal model released by Xiaomi MiMo in August 2025. It accepts text, images, and video and produces text responses for visual question answering, description, optical character understanding, and multimodal reasoning. The SFT checkpoint is primarily a starting point for additional supervised fine-tuning or reinforcement learning rather than a turnkey hosted assistant or general-purpose media-generation model.
Outputs

What MiMo-VL-7B-SFT-2508 can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
5/10 Coding
6/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiMo-VL
Model type Multimodal
Context window 128K tokens
Release date 2025-08-21
Status Current open-weight checkpoint; available for download
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for this exact checkpoint in the reviewed model card, configuration, technical report, or repository documentation.

Model notes

Canonical Hugging Face model ID: XiaomiMiMo/MiMo-VL-7B-SFT-2508. The 2508 update is a distinct checkpoint from MiMo-VL-7B-SFT and MiMo-VL-7B-RL-2508. Xiaomi describes the SFT-2508 checkpoint as a base for additional SFT and RL, while recommending the RL-2508 checkpoint for most direct usage. The model uses a Qwen2.5-VL-compatible conditional-generation architecture, has approximately 7 billion parameters, and is distributed as downloadable weights under the MIT license. The published configuration specifies a 128K maximum position length. Thinking is enabled by default; appending /no_think as the final part of a user message requests a direct response. No exact maximum output-token limit, official hosted API price, native web-search integration, prompt-caching service, batch API, or separate JSON-mode capability was verified for this checkpoint.

Cost

Model pricing

Input No official hosted API pricing; downloadable weights are available under the MIT license
Output No official hosted API pricing; downloadable weights are available under the MIT license
Model guide

MiMo-VL-7B-SFT-2508: An Open-Weight Vision-Language Checkpoint for Further Training

MiMo-VL-7B-SFT-2508 is Xiaomi MiMo’s open-weight 7-billion-parameter vision-language checkpoint for image and video understanding, multimodal reasoning, supervised fine-tuning, and reinforcement-learning research. Its 128,000-token context and MIT license make it suitable for self-hosted experimentation, but Xiaomi recommends the related RL-2508 checkpoint for most users who want to use the model directly.

What is MiMo-VL-7B-SFT-2508?

MiMo-VL-7B-SFT-2508 is Xiaomi MiMo’s supervised-fine-tuned checkpoint in the MiMo-VL 2508 model series. It is an open-weight vision-language model, meaning that its weights can be downloaded and run with suitable hardware instead of being accessed only through a provider-hosted API. The model has approximately 7 billion parameters and is distributed under the MIT license.

The “VL” designation refers to vision-language functionality: the model combines visual inputs with written instructions and returns a language response. The “SFT” designation means supervised fine-tuning. In this stage, the model has been trained on examples of desired responses so that it can follow multimodal instructions more effectively. The checkpoint is therefore useful both as a working vision-language model and as a base for researchers or developers who want to continue training it.

Xiaomi released MiMo-VL-7B-SFT-2508 on August 21, 2025. It is available as downloadable weights through Xiaomi MiMo’s Hugging Face organization and ModelScope collection. The official materials distinguish it from the earlier MiMo-VL-7B-SFT checkpoint and from the related MiMo-VL-7B-RL-2508 model.

Where it fits in Xiaomi MiMo’s lineup

The SFT-2508 checkpoint is not positioned as Xiaomi’s preferred choice for every direct-use scenario. Xiaomi describes it as a suitable base for additional supervised fine-tuning and reinforcement learning, while recommending MiMo-VL-7B-RL-2508 for most users who simply want to run the model and obtain answers.

This distinction matters in practice. A user building a research pipeline may prefer the SFT checkpoint because it provides a controlled starting point for additional training. A user looking for the best-prepared checkpoint for ordinary inference may have a better reason to consider the RL-2508 sibling. The SFT model remains the subject of this page, but its intended role is closer to an adaptable research and development checkpoint than to a managed consumer assistant.

Supported inputs and outputs

MiMo-VL-7B-SFT-2508 accepts text, images, and video. It can combine a written question with a visual input, such as asking what appears in a photograph, requesting an explanation of a chart, or asking about events in a video. Its documented visual tasks include:

  • Visual question answering
  • Image and video description
  • Optical character understanding and other image-text interpretation
  • Visual reasoning
  • Multimodal mathematical and logical problem solving

The model’s native output is text. It does not natively generate images, video, audio, speech, or music. This makes it suitable for interpreting media, not for acting as a media-generation model. It also should not be confused with a general-purpose agent: no verified native web-search, function-calling, or external tool-use capability is documented for this exact checkpoint.

Architecture and 128K context

The model combines a native-resolution vision encoder, an MLP cross-modal projector, and Xiaomi’s MiMo-7B language-model backbone. The vision encoder processes visual information, while the projector converts that information into a representation the language model can use alongside text.

The published configuration specifies a maximum position length of 128,000 tokens. A token is a small unit of text used by the model, so this limit describes the total textual context the deployment configuration can address. In multimodal use, the practical amount of material that can fit will also depend on how images or video frames are represented and on the memory available to the inference system. The 128K value is a verified configuration setting, not a guarantee that every hardware setup can process the maximum efficiently.

The model uses a Qwen2.5-VL-compatible conditional-generation architecture. Xiaomi’s repository recommends a temperature of 0.3 and top-p of 0.95 for generation. Thinking is enabled by default in the documented setup; adding /no_think as the final text in a user query requests a direct response without the visible reasoning mode.

Reasoning and coding capability

MiMo-VL-7B-SFT-2508 is designed for multimodal reasoning, including tasks that require connecting visual evidence with mathematical, logical, or textual instructions. For example, it may be used to interpret a diagram, reason about a photographed problem, or follow a sequence of events in a video. Its reasoning behavior should still be evaluated on the specific images, videos, languages, and task formats that matter to an application.

The supplied evaluation is editorial rather than a provider-published benchmark: reasoning is rated 7 out of 10, coding 5 out of 10, and speed 6 out of 10. The reasoning assessment reflects the model’s multimodal design and intended training role, not a standardized score. Coding is not its central specialization. It can produce or explain code as text, but the research does not verify a dedicated coding mode, code execution environment, software-engineering toolchain, or superior performance on general coding tasks.

Why use the SFT checkpoint?

Xiaomi describes the model’s development as a four-stage training process covering projector warmup, vision-language alignment, general multimodal pretraining, and long-context supervised fine-tuning. This training path is relevant because the checkpoint is intended to be modified or extended by users with their own data and training procedures.

Supervised fine-tuning can adapt a model to a particular response style, domain, document format, or visual task when suitable training examples are available. Reinforcement learning can then be used in some research workflows to optimize behavior against a reward or preference signal. These activities require engineering expertise, compute resources, data preparation, and evaluation; downloading the weights alone does not provide a managed training service.

The MIT license is permissive and supports a broad range of experimentation and deployment subject to the license terms. Users still need to review the license, model documentation, data rights, and any applicable rules before incorporating the model into a commercial or sensitive application.

Deployment, pricing, and API availability

There is no verified official per-token hosted API price for MiMo-VL-7B-SFT-2508. The weights are downloadable, so the direct model cost is not expressed as a provider subscription or input/output token rate. Instead, users bear the practical cost of hardware, storage, electricity, deployment, maintenance, or a third-party inference provider if they do not run the model themselves.

The model can be run with Transformers, vLLM, SGLang, or compatible local inference tools according to the supplied research. These are deployment options rather than evidence that Xiaomi operates a dedicated managed API for this exact checkpoint. No official maximum generated-token limit was verified. The maximum context configuration is known to be 128,000 tokens, but the separate maximum length of a generated answer is undocumented for this model.

Streaming support, prompt caching, batch API access, structured JSON mode, and a native web-search integration are also unverified for this exact checkpoint. Developers can potentially build surrounding infrastructure, but those additions should not be presented as native model capabilities.

Main strengths and limitations

Strengths

  • Open weights: The model can be downloaded and adapted instead of requiring access through a closed hosted endpoint.
  • Broad visual input: It supports both images and video alongside text, enabling more than ordinary text-only question answering.
  • Long context: The published 128,000-token position length is useful for long instructions and extended multimodal workflows, subject to hardware constraints.
  • Research flexibility: Xiaomi specifically presents the checkpoint as a base for further supervised fine-tuning and reinforcement learning.
  • Permissive licensing: The weights are released under the MIT license.

Limitations

  • Not a media generator: It produces text and does not natively create images, video, audio, speech, or music.
  • Not a turnkey hosted service: There is no verified Xiaomi-hosted price or dedicated API offering for this exact checkpoint.
  • Hardware responsibility: Self-hosting requires suitable compute, and the 128K context setting may be demanding in practice.
  • Unverified integrations: Native web search, tools, structured output guarantees, caching, batch inference, and a maximum output-token limit are not documented as verified features.
  • Checkpoint positioning: Xiaomi recommends the RL-2508 sibling for most users who want direct use, so the SFT version may require more customization or evaluation.

When to choose MiMo-VL-7B-SFT-2508

Choose MiMo-VL-7B-SFT-2508 when you need an open-weight model for image or video understanding and want control over deployment or continued training. It is particularly relevant for research teams testing multimodal reasoning, developers adapting a vision-language model to a specialized dataset, and organizations that prefer to operate inference on their own infrastructure rather than rely on a provider-hosted API.

Its cost profile can be attractive compared with per-token commercial services when an organization already has suitable hardware and expects substantial usage. However, “free weights” does not mean zero operating cost. Compute, engineering time, model evaluation, and infrastructure remain part of the total cost. The editorial assessment rates its cost favorably at 9 out of 10 because there is no official weight-download fee, but that is not a provider-published cost benchmark.

Another option may be more appropriate if the priority is immediate, managed access; a documented API price; guaranteed structured output; built-in web search or tool execution; or image, video, and speech generation. If the goal is ordinary direct inference rather than further training, Xiaomi’s related RL-2508 checkpoint deserves consideration because Xiaomi recommends it for most users. A text-focused coding model may also be a better fit for software development, while a dedicated generative media model is more appropriate for creating non-text content.

Bottom line

MiMo-VL-7B-SFT-2508 is best understood as an adaptable open-weight vision-language foundation checkpoint. Its combination of image and video input, text output, a 128K context configuration, MIT licensing, and support for further SFT or reinforcement-learning work makes it useful for multimodal research and self-hosted development. Its main trade-off is that it is not a fully managed product: pricing, maximum output length, tool integrations, and several production features are not verified, and Xiaomi points most direct users toward the RL-2508 alternative.


Answers to Frequently Asked Questions

Who should use MiMo-VL-7B-SFT-2508 instead of MiMo-VL-7B-RL-2508?
MiMo-VL-7B-SFT-2508 is better suited to researchers and developers who want to continue supervised fine-tuning or reinforcement learning, customize behavior, or self-host the model. Xiaomi recommends the related MiMo-VL-7B-RL-2508 checkpoint for most users who primarily want ready-to-use direct inference.
Is MiMo-VL-7B-SFT-2508 available through an official hosted API?
No verified Xiaomi-hosted API or official per-token pricing has been documented for this exact checkpoint. Users can download the weights and run them with tools such as Transformers, vLLM, or SGLang, but they are responsible for hardware, infrastructure, maintenance, or third-party inference costs.
Does MiMo-VL-7B-SFT-2508 have a 128K context window?
Its published configuration specifies a maximum position length of 128,000 tokens. The practical usable context depends on how visual inputs are represented and on available hardware, so the 128K setting does not guarantee efficient processing on every deployment.
What inputs and outputs does MiMo-VL-7B-SFT-2508 support?
The model supports text, images, and video as inputs and returns text as output. It can perform visual question answering, image and video description, optical character understanding, visual reasoning, and multimodal mathematical or logical problem solving. It does not natively generate images, video, audio, speech, or music.
What is MiMo-VL-7B-SFT-2508?
MiMo-VL-7B-SFT-2508 is Xiaomi MiMo’s approximately 7-billion-parameter, open-weight vision-language checkpoint. It accepts text, images, and video, produces text responses, is distributed under the MIT license, and is intended both for direct multimodal inference and further supervised fine-tuning or reinforcement learning.


Sources 5
Provider

About Xiaomi HyperAI