What is Kimi-VL-A3B-Thinking?
Kimi-VL-A3B-Thinking is an open-weight multimodal reasoning model from Moonshot AI, released on April 10, 2025. It belongs to the Kimi-VL family and is distributed as downloadable model weights rather than presented in the supplied research as a separately priced hosted API product. “Multimodal” means that the model can work with more than one kind of input: it accepts text and visual information, including images and documented video-understanding workloads, then returns a text response.
The model is aimed at reasoning-heavy visual tasks rather than simple description. For example, it can be used to inspect a chart and explain its trends, read text from a document image, solve a mathematics problem shown in a photograph, compare several images, or answer questions about a long visual-and-textual record. Moonshot AI trained it with long chain-of-thought supervised fine-tuning and reinforcement learning, which positions the checkpoint for problems where the answer depends on multiple reasoning steps.
Architecture and 128K context window
Kimi-VL-A3B-Thinking uses a Mixture-of-Experts, or MoE, architecture. In an MoE model, the full network contains many parameters, but only selected parts are activated for each token. The checkpoint has approximately 16 billion total parameters and approximately 3 billion activated language-model parameters. The smaller activated portion can improve computational efficiency compared with treating every parameter as active on every token, but deployment still needs to accommodate the total model and its vision components.
The model combines an MoE language model, Moonshot AI’s native-resolution MoonViT vision encoder, and an MLP projector that connects visual features to the language model. Its published context length is 131,072 tokens, commonly described as a 128K-token context window. This gives it room for long documents, extended conversations, and multiple visual inputs, although practical capacity and speed will depend on the serving stack, image or video preprocessing, available GPU memory, and the size of the generated response.
Supported inputs and outputs
| Capability | Supported status | Practical meaning |
|---|---|---|
| Text input | Yes | Prompts, questions, instructions, and long textual context |
| Image input | Yes | Visual question answering, OCR, document images, charts, and multi-image comparison |
| Video input | Documented | Video understanding and analysis, subject to the selected inference implementation |
| Text output | Yes | Answers, explanations, extracted text, and reasoning-oriented responses |
| Image, video, audio, or speech output | No | The checkpoint is not documented as a media or speech generator |
| Native computer actions | No | Agent or computer-use behavior requires an external framework |
The distinction between visual input and visual output is important. Kimi-VL-A3B-Thinking can interpret images and video, but its native output is text. It should not be selected when the requirement is image generation, video generation, speech synthesis, or another direct non-text output.
Reasoning capabilities and evaluations
The model’s primary strength is multimodal reasoning. Moonshot AI designed it for mathematical reasoning, OCR, image and video comprehension, visual question answering, long-context understanding, and visual tasks that can be incorporated into agent research. Its native-resolution visual encoder is intended to preserve useful detail in high-resolution inputs, which matters for small text, diagrams, tables, and documents.
The original technical report reports scores of 61.7 on MMMU, 36.8 on MathVision, and 71.3 on MathVista. These are release-period benchmark results from the technical report, not guarantees for every deployment or prompt. They should also not be confused with a universal “reasoning score.” The model’s reasoning quality can vary with image resolution, prompt design, decoding settings, quantization, and the inference framework used.
For coding, the supplied research supports a limited interpretation: the model can generate text and may be useful for code-related multimodal tasks, but it is not presented as a specialist coding model. Its main value is interpreting visual or document context and explaining a result, not replacing a dedicated software-engineering workflow.
Deployment, license, and tooling
The weights are available through Hugging Face under the MIT license. The first-party Kimi-VL repository documents inference routes involving Transformers, vLLM, and SGLang, as well as quantized-model tooling. These options give technical users flexibility to run the checkpoint locally or on their own infrastructure, but deployment is not lightweight simply because only about 3 billion language parameters are activated per token. The approximately 16-billion-parameter total model, vision encoder, context length, and visual preprocessing all affect memory requirements.
Kimi-VL-A3B-Thinking does not come with a documented first-party web-search tool, official hosted token pricing, batch API, structured-output guarantee, or native function-calling interface in the supplied research. Tool use and agent behavior therefore require an external serving layer or orchestration framework. The model can contribute visual reasoning to a computer-use system, but it does not itself return native operating-system actions or provide a hosted computer-control API.
Moonshot AI’s documentation recommends a temperature of 0.8 for thinking workloads. That is a provider recommendation rather than a guarantee of best results for every task. Users should test decoding settings against their own requirements, especially when extracting exact text or evaluating mathematical answers.
Pricing and current position
No official hosted input or output price is documented for this exact downloadable checkpoint in the supplied research. Its direct model cost is therefore primarily an infrastructure question: users must account for GPU hardware or rental, storage, serving software, energy, and operational maintenance. The open MIT license can be attractive for experimentation and self-hosting, but it does not make large-scale deployment cost-free.
The original checkpoint remains downloadable and usable, but it is no longer the newest model in this line. Moonshot AI identifies Kimi-VL-A3B-Thinking-2506 as an improved successor with better multimodal reasoning, general visual understanding, video performance, high-resolution input handling, and more efficient thinking-token usage. New projects should normally evaluate the 2506 release first, while the original model remains relevant when reproducibility, checkpoint compatibility, or a comparison with the earlier release matters.
Main strengths and limitations
- Long multimodal context: The 128K-token window is suitable for lengthy text combined with visual evidence, although actual throughput depends on hardware and preprocessing.
- Reasoning-oriented vision: The model targets mathematical visual reasoning, OCR, document analysis, visual question answering, and video understanding instead of only producing captions.
- Open deployment: Downloadable MIT-licensed weights and support in commonly used inference ecosystems allow local or controlled deployment.
- Parameter efficiency at inference: The MoE design activates approximately 3 billion language-model parameters per token, while the complete model contains approximately 16 billion parameters.
- Resource demands: The total model and visual components still require substantial GPU memory and engineering work, particularly for long contexts or video.
- No native media generation: It analyzes visual inputs but does not generate images, video, audio, or speech.
- No documented hosted-service guarantees: There is no supplied official price, maximum output-token limit, knowledge-cutoff date, JSON-mode guarantee, batch API, or first-party web search for this checkpoint.
- External tools are required: Browser access, function execution, computer control, and agent orchestration must be supplied by other software.
When to choose Kimi-VL-A3B-Thinking
Choose Kimi-VL-A3B-Thinking when you need an open-weight model for local research or self-hosted multimodal reasoning and you value control over the weights, deployment environment, and data path. It is a reasonable candidate for OCR experiments, image-heavy document analysis, mathematical visual question answering, multi-image comparison, and research prototypes that combine visual reasoning with an external agent framework.
It is less suitable when you need a simple managed API with published per-token pricing, guaranteed structured JSON, integrated web search, native function calling, or direct media generation. It is also a poor fit for a low-resource environment that cannot provide the memory and throughput needed by a roughly 16-billion-parameter multimodal checkpoint. For a new deployment, the 2506 successor should be tested first because Moonshot AI describes it as the improved version; the original remains the more appropriate choice when the exact earlier checkpoint is required.
Bottom line
Kimi-VL-A3B-Thinking is best understood as a self-hostable, text-output vision-language reasoning model rather than a general-purpose hosted assistant. Its defining combination is a 128K context window, native-resolution visual processing, MoE efficiency, and support for image- and video-centered reasoning tasks. Its open MIT-licensed weights make it useful for controlled experimentation, but users must provide the infrastructure and any tools around it. The absence of documented hosted pricing, structured-output guarantees, native actions, and media generation places clear boundaries around where it belongs in a production stack.

