MoonViT

MoonViT-SO-400M

by Moonshot AI · Current open-weight model; available for local use through Hugging Face Transformers; not deployed by a Hugging Face Inference Provider

MoonViT-SO-400M is Moonshot AI’s open-weight, approximately 400-million-parameter vision encoder. Initialized from SigLIP-SO-400M, it uses BF16 weights and custom Transformers code to process images at native resolution and return visual feature tensors for vision-language, document, retrieval, and computer-vision pipelines.

Embeddings
MoonViT-SO-400M is Moonshot AI’s standalone vision encoder for extracting visual features from images. Released under the MIT license through Hugging Face, it uses BF16 weights, custom Transformers code, and native-resolution image processing so that visual token counts can vary with the input image rather than being forced into one fixed square size.
Outputs

What MoonViT-SO-400M can produce

Embeddings
Inputs

What it can understand

Images
Specifications

Technical details

Model family MoonViT
Model type Other
Release date 2025-04-17
Status Current open-weight model; available for local use through Hugging Face Transformers; not deployed by a Hugging Face Inference Provider
Knowledge cutoff notes

Not applicable or publicly documented. MoonViT-SO-400M is a visual feature encoder and does not generate knowledge-bearing text.

Model notes

MoonViT-SO-400M is a vision encoder rather than a complete vision-language model. It was initialized from and continually pretrained on SigLIP-SO-400M, then released separately from Kimi-VL-A3B-Instruct. The repository contains custom Transformers code and requires trust_remote_code=True for the documented loading path. The model card reports approximately 0.4B parameters, BF16 weights, a 14-pixel patch size, 27 hidden layers, 16 attention heads, a hidden size of 1,152, and an example output tensor shape of [1092, 4, 1152]. Its processor configuration specifies a 4,096-token input limit, 1,024 pooled tokens, and native-resolution image-grid handling. No model-specific API pricing, knowledge cutoff, text context window, or maximum generated-token limit is documented.

Model guide

MoonViT-SO-400M: Moonshot AI’s Native-Resolution Vision Encoder

MoonViT-SO-400M is an open-weight, approximately 400-million-parameter native-resolution vision encoder from Moonshot AI. Initialized from SigLIP-SO-400M and continually pretrained for use in Kimi-VL and other vision-language systems, it produces image feature tensors rather than text, images, audio, or video.

What MoonViT-SO-400M is

MoonViT-SO-400M is an approximately 400-million-parameter vision encoder released by Moonshot AI. Its canonical model repository is moonshotai/MoonViT-SO-400M. A vision encoder converts an image into numerical feature tensors that another model or application can interpret. It is therefore a component for multimodal systems, not a complete conversational assistant.

The model was initialized from SigLIP-SO-400M and continually pretrained by Moonshot AI. It was developed for use in systems such as Kimi-VL, but the standalone release makes the encoder and its weights available for reuse in custom vision-language models, retrieval systems, document-processing pipelines, and other computer-vision applications.

MoonViT-SO-400M was released on April 17, 2025, according to the supplied model record. It is distributed with BF16 safetensors weights and an MIT license. The MIT license is relevant for teams considering local experimentation or integration, although the practical requirements of running the model still depend on available hardware and the downstream system connected to it.

Native-resolution architecture and image processing

The model uses a vision-transformer architecture. Its published configuration specifies 27 hidden layers, 16 attention heads, a hidden size of 1,152, an intermediate size of 4,304, and a 14-pixel patch size. Patches are small sections of an image processed as visual input units, allowing the transformer to build a representation of the image.

Its most important design distinction is native-resolution processing. Many vision encoders resize every image to one predetermined square dimension before inference. MoonViT is designed to preserve more of the source image’s resolution and aspect-ratio information. The image processor accepts image-grid information and can produce a variable number of visual tokens depending on the dimensions of the input.

This design is especially relevant to documents, screenshots, charts, and other images where resizing can make small text or fine layout details harder to preserve. It may also create a more practical representation for images with very different shapes. Native-resolution processing does not guarantee better results for every task, however: the resulting token count and resource requirements can vary with the image, and the downstream model must be designed to consume the encoder’s output.

The processor configuration specifies support for padding, a 4,096-token input limit, and 1,024 pooled tokens. These are image-processing settings rather than a conversational text context window. The supplied documentation does not provide a text context length because MoonViT-SO-400M is not a text-generating language model.

Outputs and supported modalities

MoonViT-SO-400M accepts images and returns visual feature tensors. The model card demonstrates an output with BF16 values and a shape of [1092, 4, 1152] for a sample image. The exact tensor dimensions can depend on the image and processing configuration, so this example should not be treated as a universal output shape.

The model does not independently return captions, answers, classifications, generated images, audio, or video. It also does not provide text output, function calling, web search, streaming responses, or a standalone agent interface. Those capabilities require a compatible language model, task-specific decoder, or application layer connected to the visual features.

CapabilityMoonViT-SO-400M
Image inputYes
Text inputNot documented as a model input
Visual feature outputYes
Text generationNo
Image, audio, or video generationNo
Standalone chat or reasoningNo
Tool or function useNo

How it is deployed

The documented loading path uses Hugging Face Transformers with AutoModel and AutoImageProcessor. The repository supplies custom configuration, preprocessing, and modeling implementations, so the documented setup requires trust_remote_code=True. Developers should review that dependency before using the model in a security-sensitive or tightly controlled environment.

A typical implementation loads the image processor to prepare an image, loads the model, and then passes the processed image into the encoder to obtain visual features. The features can subsequently be passed to a language-model connector, a retrieval index, a classifier, or another downstream component. MoonViT-SO-400M is therefore best understood as one layer in a larger pipeline rather than an end-user application.

The model is available for local use through its Hugging Face repository. The supplied research states that it is not currently deployed through a Hugging Face Inference Provider. No model-specific managed API endpoint, hosted inference tier, or official token-based service is documented.

Strengths and practical use cases

The main practical strength is flexible image-resolution handling. For high-resolution documents, diagrams, screenshots, and images with non-square layouts, retaining more spatial information can be more useful than applying one fixed resize to every input. The model’s open-weight distribution and MIT license also make it suitable for researchers and developers who want to assemble a custom multimodal stack.

  • Vision-language backbones: Connect its visual features to a language model for image questions, captioning, document analysis, or other multimodal tasks.
  • High-resolution document understanding: Use it as the image-processing component for pages, forms, charts, or screenshots where layout and small visual details matter.
  • Image feature extraction: Produce reusable representations for downstream computer-vision experiments.
  • Image retrieval and similarity: Feed visual embeddings or feature representations into a search or matching system, provided the downstream application is designed for the model’s output.
  • Multimodal research: Experiment with connectors, decoders, and custom vision-language architectures without relying on a hosted endpoint.

The supplied research does not report standalone benchmark scores for MoonViT-SO-400M. Results or capabilities associated with Kimi-VL should not automatically be presented as results for this encoder by itself.

Limitations and missing capabilities

MoonViT-SO-400M cannot be used as a drop-in replacement for a chat model. It does not generate text, answer questions independently, perform general reasoning, write code, browse the web, call tools, or manage an agent workflow. A separate language model and an appropriate connector are required for these tasks.

Its documented input limit is expressed through the image processor’s 4,096-token setting, not through a conventional text context window. The research does not specify a maximum generated-token limit because the model generates no text. It also does not publish model-specific API prices, text-token pricing, a knowledge cutoff, or a managed production SLA.

Local deployment introduces responsibilities that a hosted model would normally handle for the user. Teams must provide suitable compute, manage model files, verify preprocessing compatibility, and build or select the downstream projector or decoder. Variable-resolution processing may also require more careful memory planning than a fixed-size image pipeline, especially when processing large images or batches.

The model’s documentation does not establish standalone reasoning, coding, speed, or cost scores. Any evaluation of those properties is therefore an application-level judgment rather than a provider-published rating. In practice, total cost and latency will depend on hardware, image size, batching, precision, and the additional model used after the encoder.

Pricing and availability

MoonViT-SO-400M has no documented provider-hosted API price in the supplied research. The model is available as open weights through Hugging Face, so there is no listed per-input or per-output token charge for the local model itself. Local use is not necessarily free in operational terms: hardware, storage, inference, engineering, and hosting costs still apply.

The model is not listed as being served by a Hugging Face Inference Provider. Developers should not assume that a public managed endpoint, service-level guarantee, or production-ready API is available simply because the repository is public.

When to choose MoonViT-SO-400M

Choose MoonViT-SO-400M when you need an open-weight visual encoder and want direct control over the image-processing and multimodal pipeline. It is a reasonable candidate for native-resolution document or image experiments, custom vision-language backbones, and applications where local deployment or the MIT license is more important than turnkey access.

It is less appropriate when you need an immediately usable assistant, image question answering without assembling additional components, predictable hosted pricing, or a managed API. A complete vision-language model is a better fit for those requirements. A fixed-resolution encoder may also be preferable when simple, predictable memory usage and uniform batching matter more than preserving variable image resolution.

MoonViT-SO-400M should also be distinguished from the broader Kimi product ecosystem. Moonshot AI’s consumer and multimodal products may provide conversation, research, tools, or creative features, but those capabilities should not be attributed to this standalone encoder. MoonViT’s role is narrower: it transforms images into visual representations that another system can use.

Bottom line

MoonViT-SO-400M is a specialized, open-weight vision component rather than a general-purpose AI model. Its defining feature is native-resolution image processing, supported by a roughly 400-million-parameter transformer with a 14-pixel patch size and variable visual-token output. That makes it relevant to developers building high-resolution vision-language or retrieval systems, particularly when local control is valuable.

The trade-off is integration work. There is no standalone text output, conversational interface, hosted pricing, or documented provider inference tier. Its value depends on the quality of the downstream language model or task-specific system connected to it.


Answers to Frequently Asked Questions

Is MoonViT-SO-400M available through a hosted API, and what does it cost?
The supplied documentation does not list a provider-hosted API, managed inference tier, or model-specific API pricing for MoonViT-SO-400M. Its BF16 open weights are available through Hugging Face under an MIT license, but local deployment still involves hardware, storage, inference, engineering, and hosting costs.
How can developers deploy MoonViT-SO-400M?
Developers can load MoonViT-SO-400M locally from its Hugging Face repository using Transformers with AutoModel and AutoImageProcessor. The documented setup requires trust_remote_code=True because the repository provides custom implementations. The resulting visual features can be connected to a language model, classifier, retrieval index, or another downstream system.
Can MoonViT-SO-400M generate text or answer questions by itself?
No. MoonViT-SO-400M outputs visual feature tensors but does not independently generate captions, answer questions, perform chat, write code, browse the web, or use tools. These capabilities require a compatible language model, connector, decoder, or application layer.
What does native-resolution processing mean in MoonViT-SO-400M?
Native-resolution processing allows MoonViT-SO-400M to preserve more of an image’s original resolution and aspect-ratio information instead of resizing every image to one fixed square size. This can be useful for documents, screenshots, charts, and images containing small text or detailed layouts. Token counts and resource requirements vary with the input image.
What is MoonViT-SO-400M?
MoonViT-SO-400M is an approximately 400-million-parameter vision encoder released by Moonshot AI. It converts images into visual feature tensors for use in multimodal systems, retrieval applications, document-processing pipelines, and other computer-vision tasks. It is not a standalone conversational or text-generation model.


Sources 5
Provider

About Moonshot AI