What MoonViT-SO-400M is
MoonViT-SO-400M is an approximately 400-million-parameter vision encoder released by Moonshot AI. Its canonical model repository is moonshotai/MoonViT-SO-400M. A vision encoder converts an image into numerical feature tensors that another model or application can interpret. It is therefore a component for multimodal systems, not a complete conversational assistant.
The model was initialized from SigLIP-SO-400M and continually pretrained by Moonshot AI. It was developed for use in systems such as Kimi-VL, but the standalone release makes the encoder and its weights available for reuse in custom vision-language models, retrieval systems, document-processing pipelines, and other computer-vision applications.
MoonViT-SO-400M was released on April 17, 2025, according to the supplied model record. It is distributed with BF16 safetensors weights and an MIT license. The MIT license is relevant for teams considering local experimentation or integration, although the practical requirements of running the model still depend on available hardware and the downstream system connected to it.
Native-resolution architecture and image processing
The model uses a vision-transformer architecture. Its published configuration specifies 27 hidden layers, 16 attention heads, a hidden size of 1,152, an intermediate size of 4,304, and a 14-pixel patch size. Patches are small sections of an image processed as visual input units, allowing the transformer to build a representation of the image.
Its most important design distinction is native-resolution processing. Many vision encoders resize every image to one predetermined square dimension before inference. MoonViT is designed to preserve more of the source image’s resolution and aspect-ratio information. The image processor accepts image-grid information and can produce a variable number of visual tokens depending on the dimensions of the input.
This design is especially relevant to documents, screenshots, charts, and other images where resizing can make small text or fine layout details harder to preserve. It may also create a more practical representation for images with very different shapes. Native-resolution processing does not guarantee better results for every task, however: the resulting token count and resource requirements can vary with the image, and the downstream model must be designed to consume the encoder’s output.
The processor configuration specifies support for padding, a 4,096-token input limit, and 1,024 pooled tokens. These are image-processing settings rather than a conversational text context window. The supplied documentation does not provide a text context length because MoonViT-SO-400M is not a text-generating language model.
Outputs and supported modalities
MoonViT-SO-400M accepts images and returns visual feature tensors. The model card demonstrates an output with BF16 values and a shape of [1092, 4, 1152] for a sample image. The exact tensor dimensions can depend on the image and processing configuration, so this example should not be treated as a universal output shape.
The model does not independently return captions, answers, classifications, generated images, audio, or video. It also does not provide text output, function calling, web search, streaming responses, or a standalone agent interface. Those capabilities require a compatible language model, task-specific decoder, or application layer connected to the visual features.
| Capability | MoonViT-SO-400M |
|---|---|
| Image input | Yes |
| Text input | Not documented as a model input |
| Visual feature output | Yes |
| Text generation | No |
| Image, audio, or video generation | No |
| Standalone chat or reasoning | No |
| Tool or function use | No |
How it is deployed
The documented loading path uses Hugging Face Transformers with AutoModel and AutoImageProcessor. The repository supplies custom configuration, preprocessing, and modeling implementations, so the documented setup requires trust_remote_code=True. Developers should review that dependency before using the model in a security-sensitive or tightly controlled environment.
A typical implementation loads the image processor to prepare an image, loads the model, and then passes the processed image into the encoder to obtain visual features. The features can subsequently be passed to a language-model connector, a retrieval index, a classifier, or another downstream component. MoonViT-SO-400M is therefore best understood as one layer in a larger pipeline rather than an end-user application.
The model is available for local use through its Hugging Face repository. The supplied research states that it is not currently deployed through a Hugging Face Inference Provider. No model-specific managed API endpoint, hosted inference tier, or official token-based service is documented.
Strengths and practical use cases
The main practical strength is flexible image-resolution handling. For high-resolution documents, diagrams, screenshots, and images with non-square layouts, retaining more spatial information can be more useful than applying one fixed resize to every input. The model’s open-weight distribution and MIT license also make it suitable for researchers and developers who want to assemble a custom multimodal stack.
- Vision-language backbones: Connect its visual features to a language model for image questions, captioning, document analysis, or other multimodal tasks.
- High-resolution document understanding: Use it as the image-processing component for pages, forms, charts, or screenshots where layout and small visual details matter.
- Image feature extraction: Produce reusable representations for downstream computer-vision experiments.
- Image retrieval and similarity: Feed visual embeddings or feature representations into a search or matching system, provided the downstream application is designed for the model’s output.
- Multimodal research: Experiment with connectors, decoders, and custom vision-language architectures without relying on a hosted endpoint.
The supplied research does not report standalone benchmark scores for MoonViT-SO-400M. Results or capabilities associated with Kimi-VL should not automatically be presented as results for this encoder by itself.
Limitations and missing capabilities
MoonViT-SO-400M cannot be used as a drop-in replacement for a chat model. It does not generate text, answer questions independently, perform general reasoning, write code, browse the web, call tools, or manage an agent workflow. A separate language model and an appropriate connector are required for these tasks.
Its documented input limit is expressed through the image processor’s 4,096-token setting, not through a conventional text context window. The research does not specify a maximum generated-token limit because the model generates no text. It also does not publish model-specific API prices, text-token pricing, a knowledge cutoff, or a managed production SLA.
Local deployment introduces responsibilities that a hosted model would normally handle for the user. Teams must provide suitable compute, manage model files, verify preprocessing compatibility, and build or select the downstream projector or decoder. Variable-resolution processing may also require more careful memory planning than a fixed-size image pipeline, especially when processing large images or batches.
The model’s documentation does not establish standalone reasoning, coding, speed, or cost scores. Any evaluation of those properties is therefore an application-level judgment rather than a provider-published rating. In practice, total cost and latency will depend on hardware, image size, batching, precision, and the additional model used after the encoder.
Pricing and availability
MoonViT-SO-400M has no documented provider-hosted API price in the supplied research. The model is available as open weights through Hugging Face, so there is no listed per-input or per-output token charge for the local model itself. Local use is not necessarily free in operational terms: hardware, storage, inference, engineering, and hosting costs still apply.
The model is not listed as being served by a Hugging Face Inference Provider. Developers should not assume that a public managed endpoint, service-level guarantee, or production-ready API is available simply because the repository is public.
When to choose MoonViT-SO-400M
Choose MoonViT-SO-400M when you need an open-weight visual encoder and want direct control over the image-processing and multimodal pipeline. It is a reasonable candidate for native-resolution document or image experiments, custom vision-language backbones, and applications where local deployment or the MIT license is more important than turnkey access.
It is less appropriate when you need an immediately usable assistant, image question answering without assembling additional components, predictable hosted pricing, or a managed API. A complete vision-language model is a better fit for those requirements. A fixed-resolution encoder may also be preferable when simple, predictable memory usage and uniform batching matter more than preserving variable image resolution.
MoonViT-SO-400M should also be distinguished from the broader Kimi product ecosystem. Moonshot AI’s consumer and multimodal products may provide conversation, research, tools, or creative features, but those capabilities should not be attributed to this standalone encoder. MoonViT’s role is narrower: it transforms images into visual representations that another system can use.
Bottom line
MoonViT-SO-400M is a specialized, open-weight vision component rather than a general-purpose AI model. Its defining feature is native-resolution image processing, supported by a roughly 400-million-parameter transformer with a 14-pixel patch size and variable visual-token output. That makes it relevant to developers building high-resolution vision-language or retrieval systems, particularly when local control is valuable.
The trade-off is integration work. There is no standalone text output, conversational interface, hosted pricing, or documented provider inference tier. Its value depends on the quality of the downstream language model or task-specific system connected to it.

