Kimi-VL

Kimi-VL-A3B-Instruct

by Moonshot AI · Current open-weight/downloadable model; a newer Kimi-VL-A3B-Thinking-2506 variant is recommended for stronger multimodal reasoning

Moonshot AI’s Kimi-VL-A3B-Instruct is an MIT-licensed open-weight vision-language model for image, video, OCR, document, and screenshot understanding. It combines a native-resolution vision encoder with a mixture-of-experts decoder, contains approximately 16B total parameters with about 3B activated per step, and supports a 128K-token context window for self-hosted multimodal applications.

Text Reasoning Coding
Kimi-VL-A3B-Instruct is an open-weight multimodal model from Moonshot AI built to understand visual and textual information rather than generate images, video, audio, or speech. It can process single or multiple images, video, long documents, OCR-heavy material, and computer-interface screenshots, then produce text responses. Its combination of a native-resolution vision encoder, mixture-of-experts language decoder, and 128K context window makes it particularly relevant for self-hosted multimodal applications that need broad visual input without activating the full parameter count on every inference step.
Outputs

What Kimi-VL-A3B-Instruct can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
6/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Kimi-VL
Model type Multimodal
Context window 131K tokens
Release date 2025-04-10
Status Current open-weight/downloadable model; a newer Kimi-VL-A3B-Thinking-2506 variant is recommended for stronger multimodal reasoning
Knowledge cutoff notes

No authoritative exact knowledge-cutoff date was identified in the model card, technical report, or project documentation.

Model notes

Open-weight MIT-licensed model identified by the Hugging Face repository moonshotai/Kimi-VL-A3B-Instruct. The architecture combines a native-resolution MoonViT vision encoder with a mixture-of-experts language decoder. The model has approximately 16B total parameters and activates approximately 3B language parameters. Official project documentation lists a 128K context length and recommends the Instruct variant for general multimodal perception, OCR, long-video and long-document understanding, video perception, and OS-agent use. Deployment examples use Transformers, vLLM, and SGLang with trust_remote_code. Fine-tuning is documented through LLaMA-Factory. Tool execution, structured-output guarantees, batching, caching, and hosted pricing are not established as intrinsic features of this exact checkpoint.

Cost

Model pricing

Input No official Moonshot-hosted API price published; self-hosted/open-weight model
Output No official Moonshot-hosted API price published; self-hosted/open-weight model
Model guide

Kimi-VL-A3B-Instruct: Efficient Open-Weight Vision and Video Understanding

Kimi-VL-A3B-Instruct is Moonshot AI’s MIT-licensed, open-weight vision-language model for image, video, OCR, document, and screenshot understanding. Its mixture-of-experts architecture contains approximately 16 billion total parameters but activates about 3 billion language parameters per inference step, and it supports a 128K-token context window.

What Kimi-VL-A3B-Instruct is

Kimi-VL-A3B-Instruct is Moonshot AI’s instruction-tuned, open-weight vision-language model. In practical terms, it accepts text together with visual inputs and returns text that describes, analyzes, compares, or reasons about those inputs. Typical tasks include asking questions about an image, extracting text from a scanned page, comparing several images, examining a video, or interpreting a screenshot of a software interface.

The model is intended for multimodal understanding, not media creation. It does not natively produce images, video, audio, speech, or music. This distinction is important because Moonshot AI’s broader Kimi ecosystem includes products and plugins with creative-generation features, but those capabilities should not be attributed to this downloadable checkpoint.

Kimi-VL-A3B-Instruct was released on April 10, 2025, according to the supplied technical-report information. The weights are available through the Moonshot AI project and its Hugging Face model repository under the MIT license, allowing local deployment and modification subject to that license.

Architecture and 128K context window

The model combines three main components: a MoonViT vision encoder, an MLP projector, and a mixture-of-experts language decoder. The vision encoder uses native-resolution processing, which is designed to retain useful detail across images with different sizes and aspect ratios rather than treating every image as an identically sized input.

Its language component has approximately 16 billion total parameters, while approximately 3 billion language-model parameters are activated for each inference step. A mixture-of-experts model contains several specialized parameter groups but routes each token through only part of the network. This can reduce the amount of computation required for an individual request compared with running all parameters of an equivalent dense model. It does not mean that the model requires only 3 billion parameters to download or that every deployment will have the same speed; memory use, visual-token processing, quantization, batching, and hardware still matter.

The documented context length is 128K tokens. A large context makes it possible to combine long text with visual material, review extended documents, compare multiple images, or support long-video workflows when the serving configuration and available hardware can handle the required context. The supplied research does not establish a separate maximum output-token limit, so no fixed maximum should be assumed.

Supported inputs and outputs

CapabilityStatusPractical meaning
Text inputSupportedUse written instructions or questions with multimodal prompts.
Image inputSupportedAnalyze one image or compare multiple images.
Video inputSupportedPerform video perception and understanding tasks.
Audio inputNot establishedThe supplied model information does not identify native audio understanding.
Text outputSupportedReturn explanations, answers, extracted text, and other generated text.
Image, video, or audio outputNot supported as native model outputThe checkpoint is an understanding and text-generation model.

The model is suitable for visual question answering, image-text reasoning, multi-image comparison, OCR, document comprehension, video perception, and screenshot interpretation. It can also support visual computer-use or operating-system-agent systems as a perception component. However, perception is not the same as autonomous control: browser actions, keyboard and mouse execution, tool calls, safety checks, and application state management must be supplied by the surrounding software.

Where it fits in the Kimi-VL family

Kimi-VL-A3B-Instruct is positioned for general multimodal perception and instruction following. The project documentation recommends it for OCR, long-video and long-document understanding, video perception, and OS-agent-oriented use cases. It is the practical choice in the supplied Kimi-VL information when the application needs a general-purpose visual model that can be deployed locally.

Moonshot AI also identifies Kimi-VL-A3B-Thinking-2506 as a separate family variant for more demanding long-form multimodal reasoning. That comparison does not make the Instruct model unsuitable for reasoning; rather, it indicates that the Thinking variant is the more appropriate option when extended deliberate reasoning is the main requirement. The current model remains better described as a general multimodal perception and instruction model.

Reasoning, coding, and tool use

Kimi-VL-A3B-Instruct can reason about the content of images, videos, documents, and screenshots, including relationships between visual evidence and written instructions. Its model documentation supports use for visual question answering, OCR, long-context analysis, and agent-oriented perception. The supplied research does not provide a standardized reasoning benchmark or a guaranteed reasoning mode, so its reasoning ability should be evaluated against the specific visual tasks in an application.

It can generate text and may help interpret code shown in screenshots or documents, but it is not primarily presented as a coding model. The research gives it a moderate editorial coding assessment rather than a provider-published coding score. That assessment should not be treated as a benchmark result or a guarantee of software-engineering performance.

Tool use is not an intrinsic capability established for this checkpoint. A developer can connect its text responses to OCR pipelines, browsers, computer-use controllers, or other tools, but the model itself does not supply guaranteed function calling, browser execution, or action outputs. Structured JSON or JSON-schema compliance is also not established. Applications that require reliable machine-readable output should add validation and error handling around the model response.

Deployment and fine-tuning

The official project documentation covers Transformers, vLLM, and SGLang deployment. When an inference engine such as vLLM or SGLang is configured to expose an OpenAI-compatible endpoint, applications can communicate with the served model through that endpoint. This compatibility belongs to the serving layer; it should not be confused with a Moonshot-hosted API for this exact model.

Transformers examples require custom model code through the trust_remote_code option. Operators should therefore review the code and deployment environment before enabling remote repository code in production. vLLM configurations can expose chat-completions-compatible serving and allow operators to set context-length and image-related limits, but the appropriate values depend on available hardware and the chosen serving configuration.

Fine-tuning is documented through community tooling including LLaMA-Factory. The project information describes LoRA fine-tuning on a single GPU with approximately 50 GB of VRAM, as well as multi-GPU full or LoRA fine-tuning with DeepSpeed ZeRO-2. These are deployment and training guidance points rather than a universal hardware guarantee; memory requirements can change with sequence length, image resolution, batch size, precision, and checkpoint settings.

Pricing and operating cost

There is no official Moonshot-hosted per-token API price established for Kimi-VL-A3B-Instruct in the supplied research. The model is open-weight and can be downloaded for self-hosted use under the MIT license, so the direct model license cost is not the same as a recurring hosted-model subscription.

Self-hosting still has operating costs. Users may need suitable GPU memory, storage, electricity, cloud GPU rental, engineering time, and an inference server. The mixture-of-experts design may offer a favorable compute trade-off because approximately 3 billion language parameters are activated per step, but total model size, multimodal preprocessing, long contexts, video inputs, and batching affect actual cost. No reliable universal cost-per-request figure can be calculated from the supplied information.

Main strengths and limitations

Strengths

  • Broad visual coverage: It handles image, multi-image, video, document, OCR, and screenshot-understanding tasks.
  • Long context: The documented 128K-token window is useful for long documents and extended multimodal prompts.
  • Efficient architecture: Approximately 3B language parameters are activated from an approximately 16B-parameter mixture-of-experts model for each inference step.
  • Self-hosting: Open weights, an MIT license, and documented Transformers, vLLM, and SGLang paths give developers control over deployment.
  • Fine-tuning options: The project documents LoRA and multi-GPU fine-tuning routes through available tooling.

Limitations

  • No native media generation: It produces text and does not generate images, video, audio, speech, or music.
  • No established hosted pricing: Users should not assume that this checkpoint is available through Moonshot AI’s standard paid API catalog.
  • Deployment responsibility: Hardware selection, serving, scaling, monitoring, batching, and security are the operator’s responsibility when self-hosting.
  • No guaranteed structured output: Function calling, JSON-schema output, and action execution are not established intrinsic features.
  • Reasoning trade-off: The Instruct variant is intended for general perception; the Kimi-VL-A3B-Thinking-2506 variant may be more appropriate for demanding long-form multimodal reasoning.

When to choose Kimi-VL-A3B-Instruct

Choose Kimi-VL-A3B-Instruct when you need a self-hosted model that can inspect images, video, scans, documents, or screenshots and return text-based analysis. It is a particularly reasonable fit for private OCR services, visual document workflows, image and video question answering, screenshot analysis, research tools, and perception modules inside larger computer-use systems.

Its open-weight availability is also valuable when a project needs control over deployment, data handling, or model customization rather than a fully managed API. The long context window is useful for applications that combine substantial written material with visual evidence, although the hardware cost of using the full context should be tested rather than assumed.

Consider another option when the application needs image or video generation, speech or audio processing, guaranteed JSON-schema responses, built-in tool execution, or a managed commercial endpoint with published per-token pricing. For higher-end deliberate multimodal reasoning within the same family, the supplied research points to Kimi-VL-A3B-Thinking-2506 as the more suitable alternative. For routine visual understanding where local operation and cost control matter, Kimi-VL-A3B-Instruct offers a clearer balance between broad input support and relatively low activated language-model computation.


Answers to Frequently Asked Questions

Does Kimi-VL-A3B-Instruct support tool use, function calling, or structured JSON output?
These capabilities are not established as intrinsic features of the checkpoint. Developers can connect the model to browsers, OCR systems, computer-use controllers, or other tools, but reliable function calling, JSON-schema compliance, and action execution must be provided and validated by surrounding software.
Can Kimi-VL-A3B-Instruct be deployed locally?
Yes. The model is available as open weights under the MIT license and can be deployed locally using Transformers, vLLM, or SGLang. Self-hosting still requires suitable hardware, storage, inference infrastructure, monitoring, and security controls.
How large is Kimi-VL-A3B-Instruct’s context window?
Kimi-VL-A3B-Instruct has a documented 128K-token context window. This supports long documents, extended multimodal prompts, multi-image comparisons, and long-video workflows when the serving configuration and hardware can handle the required memory and processing.
What is Kimi-VL-A3B-Instruct?
Kimi-VL-A3B-Instruct is Moonshot AI’s open-weight, instruction-tuned vision-language model. It accepts text, images, and videos and generates text for tasks such as visual question answering, OCR, document analysis, screenshot interpretation, and video understanding.
What inputs and outputs does Kimi-VL-A3B-Instruct support?
The model supports text, image, and video inputs and produces text outputs. It does not natively generate images, video, audio, speech, or music, and native audio understanding is not established in the supplied documentation.


Sources 4
Provider

About Moonshot AI