Moonlight

Moonlight-16B-A3B

by Moonshot AI · Available open-weight pretrained checkpoint

Moonlight-16B-A3B is Moonshot AI’s open-weight pretrained mixture-of-experts language model. It has approximately 16B total parameters, activates about 3B per token, supports text-only generation with an 8K context, and is licensed under MIT. The model is intended for self-hosted inference, research, coding and mathematics experimentation, and fine-tuning rather than turnkey chat or hosted API use.

Text Reasoning Coding
Moonlight-16B-A3B is Moonshot AI’s base language model for users who want open weights rather than a turnkey hosted assistant or paid API. Its sparse mixture-of-experts design combines a 16-billion-parameter model with roughly 3 billion active parameters per token. The result is a research-oriented text model that can be run with Transformers, vLLM, SGLang, and compatible infrastructure, provided the deployment environment can store the full checkpoint.
Outputs

What Moonlight-16B-A3B can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

6/10 Reasoning
6/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Moonlight
Model type General Purpose
Context window 8K tokens
Release date 2025-02-24
Status Available open-weight pretrained checkpoint
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was identified in the official model card or technical report.

Model notes

This is the pretrained base checkpoint, not the separately released Moonlight-16B-A3B-Instruct model. The model has approximately 16B total parameters and approximately 3B active parameters per token, uses a DeepSeek-V3-style MoE architecture, and was trained on 5.7T tokens with the Muon optimizer. The official model card documents Transformers, vLLM, and SGLang deployment. No official first-party hosted token pricing was found because the checkpoint is distributed as open weights. The published 8K context length is verified; a separate exact maximum output-token limit was not specified.

Model guide

Moonlight-16B-A3B: Efficient Open-Weight MoE Model for Self-Hosted Text Generation

Moonlight-16B-A3B is an open-weight, pretrained mixture-of-experts language model from Moonshot AI. It contains approximately 16 billion total parameters but activates about 3 billion per token, providing a potentially more efficient alternative to comparably large dense models. Trained on 5.7 trillion tokens with the Muon optimizer, it is designed for self-hosted text generation, language-model research, coding and mathematics experimentation, and further fine-tuning.

What is Moonlight-16B-A3B?

Moonlight-16B-A3B is an open-weight pretrained language model developed by Moonshot AI. It is intended primarily for text generation and model research rather than direct use as a consumer chatbot. The checkpoint is available through Hugging Face as moonshotai/Moonlight-16B-A3B and is listed under the MIT license.

The name describes its approximate scale: the model has 16 billion total parameters and activates about 3 billion parameters for each token. Parameters are the learned numerical values used by a language model to process and generate text. Because a mixture-of-experts model activates only selected parts of its network for each token, it can offer lower per-token computation than a dense model with a similar total parameter count. This does not eliminate the need to store the complete model, however.

Moonlight-16B-A3B is the base checkpoint in the Moonlight family. Moonshot AI also released Moonlight-16B-A3B-Instruct, an instruction-tuned variant. The two should not be treated as interchangeable: the model covered here is pretrained and may need additional prompting, fine-tuning, or post-processing for reliable assistant-style interactions.

Architecture and training approach

Moonlight-16B-A3B uses a DeepSeek-V3-style mixture-of-experts architecture. In practical terms, the model contains multiple expert components, while a routing mechanism selects which components contribute to each token. Approximately 3 billion parameters are active at a time even though the full checkpoint contains about 16 billion parameters.

Moonshot AI reports that the model was trained on 5.7 trillion tokens. Its development is associated with research into scaling the Muon optimizer to large language models. The technical work discusses weight decay, consistent update scaling, and distributed optimization methods intended to improve training efficiency. These details describe the training research behind the release; they should not be interpreted as a guarantee of performance for every downstream task.

The published model information identifies an 8K context length. A context window is the amount of text the model can consider in one request, including the prompt and the generated continuation. The supplied specifications do not identify a separate maximum output-token limit, so deployments should not assume one beyond the limits imposed by the model and serving system.

Capabilities and supported modalities

Moonlight-16B-A3B is a text-only model. It accepts text input and produces text output; native image, audio, and video input or output are not documented. It is therefore suited to conventional language-model workloads such as completion, transformation, summarization, code generation, mathematical experimentation, and research into language-model behavior.

Its base-model status is important when evaluating those capabilities. A pretrained model learns to continue text but is not necessarily optimized to follow natural-language instructions in the manner of a chat assistant. It may respond better to carefully designed completion prompts, demonstrations, task-specific fine-tuning, or an application layer that validates and formats its output.

Moonshot AI’s technical report includes evaluations covering English and Chinese understanding, reasoning, mathematics, and code. The reported benchmark areas include MMLU, MMLU-Pro, BBH, TriviaQA, HumanEval, MBPP, GSM8K, MATH, CMath, C-Eval, and CMMLU. These are release-era research results rather than guarantees for a particular prompt, language mix, quantization method, or production workload.

Coding, reasoning, and tool use

Moonlight-16B-A3B can be used for code generation and mathematical tasks, and the release materials include coding and reasoning evaluations. Its useful role is best understood as a general-purpose text model that can be adapted for these workloads, not as a specialized coding agent or reasoning system with built-in tools.

No native web search, browsing, function calling, or external action capability is specified for this checkpoint. A developer can place the model inside a larger application that supplies tools, retrieves documents, executes code, or validates responses, but those features would come from the surrounding system rather than from the model itself. Similarly, structured output or JSON behavior should be implemented and checked by the application unless the selected serving stack adds an appropriate constrained-decoding feature.

For reasoning-heavy tasks, the model can generate explanations or intermediate text, but the supplied research does not identify a dedicated reasoning mode or a provider-defined reasoning control. Results should therefore be tested on the actual task, especially when mathematical correctness or executable code matters.

Deployment, speed, and cost trade-offs

The official model card documents deployment with Transformers, vLLM, and SGLang. These tools can support local or hosted inference, and vLLM and SGLang can expose an OpenAI-compatible local HTTP interface. The exact command, memory requirement, throughput, and latency depend on hardware, precision, batching, sequence length, and serving configuration.

The approximately 3-billion active-parameter figure helps explain the model’s efficiency appeal, but it should not be confused with a 3B checkpoint. The full approximately 16B model must still be stored or made available to the inference engine. The original release is available in BF16, and community quantized versions may reduce memory usage. Quantization can change speed, quality, and compatibility, so results from a quantized community build should be evaluated separately from the original checkpoint.

Moonlight-16B-A3B has no official first-party hosted token price in the supplied research. It is distributed as open weights, meaning that users generally pay for their own hardware, hosted GPU time, storage, networking, and operations rather than for tokens supplied by Moonshot AI. This can be economical for sustained workloads, experimentation, or privacy-sensitive deployments, but it also shifts infrastructure and maintenance responsibilities to the user.

Main strengths

  • Sparse architecture: roughly 3B parameters are active per token within a 16B-parameter model, creating a potentially favorable compute profile for self-hosted inference.
  • Open deployment: the checkpoint can be downloaded and used with documented open tooling rather than requiring a Moonshot AI-hosted endpoint.
  • Broad text research scope: the model is relevant to language understanding, generation, coding, mathematics, and English-Chinese evaluation work.
  • Permissive licensing: the model card lists the MIT license, subject to the license terms and other applicable legal or operational obligations.
  • Large training exposure: Moonshot AI reports pretraining on 5.7 trillion tokens, providing substantial scale for a research checkpoint.

Limitations to plan for

  • Not instruction-tuned: the base checkpoint may not reliably behave like a polished conversational assistant. The separate Instruct model is more directly relevant when instruction following is the priority.
  • Only an 8K context: it is less suitable for very long documents or workflows that require retaining large amounts of conversation and source material in one request.
  • Full-model storage remains necessary: sparse activation lowers active computation but does not turn the model into a small 3B-weight checkpoint.
  • No documented native multimodality: image, audio, and video tasks require another model or an external processing pipeline.
  • No first-party hosted pricing or turnkey API is specified: users must select and operate compatible inference infrastructure.
  • Benchmark results are not universal guarantees: actual quality varies with prompting, language, hardware, precision, and application design.

When to choose Moonlight-16B-A3B

Choose Moonlight-16B-A3B when you need an open-weight text model that can be deployed under your own control and you are prepared to manage inference infrastructure. It is a good candidate for language-model experimentation, self-hosted completion, code and mathematics research, fine-tuning investigations, and applications where an active-parameter-efficient MoE design is more attractive than a similarly sized dense model.

It may also fit teams that want to inspect or modify the model rather than depend on a provider-managed API. The MIT license and compatibility with common serving systems can simplify experimentation, although production users still need to review licensing, security, model quality, and operational requirements.

Another option may be more appropriate when the priority is reliable instruction following, a fully managed API, native multimodal processing, built-in tools, very long context, or a predictable per-token service price. Within the same family, Moonlight-16B-A3B-Instruct is the more relevant comparison for assistant-style interactions, while the base model remains the better fit for studying or adapting a pretrained checkpoint.

Bottom line

Moonlight-16B-A3B is a research-oriented open language model whose main distinction is the combination of a 16B total-parameter mixture-of-experts architecture and approximately 3B active parameters per token. It offers a practical starting point for self-hosted text generation and model research, with documented support for Transformers, vLLM, and SGLang. Its trade-offs are equally clear: it is text-only, limited to an 8K context, not instruction-tuned, and not supplied as a priced hosted API. Those characteristics make it most useful to technically capable users who value deployment control and openness over turnkey assistant features.


Answers to Frequently Asked Questions

What are the main limitations of Moonlight-16B-A3B?
Its main limitations are the 8K context window, lack of instruction tuning, text-only operation, need to store the full 16B-parameter model despite sparse activation, and absence of an official first-party hosted token price or turnkey API. Benchmark results may also vary with prompting, language, precision, quantization, and deployment setup.
Does Moonlight-16B-A3B support multimodal input, web browsing, or tool calling?
Moonlight-16B-A3B is documented as a text-only model and has no native image, audio, video, web browsing, function-calling, or external-action capabilities. These features can be added by an application layer that connects the model to retrieval systems, tools, code execution, or constrained output validation.
What hardware and deployment options support Moonlight-16B-A3B?
The model can be deployed with Transformers, vLLM, and SGLang. Although only about 3 billion parameters are active for each token, the complete approximately 16-billion-parameter checkpoint still needs to be stored or loaded by the inference system. Hardware requirements depend on precision, quantization, context length, batching, and serving configuration.
What is Moonlight-16B-A3B?
Moonlight-16B-A3B is an open-weight, pretrained language model from Moonshot AI designed for text generation, self-hosted inference, and model research. It has approximately 16 billion total parameters and activates about 3 billion parameters per token.
Is Moonlight-16B-A3B an instruction-tuned chatbot model?
No. Moonlight-16B-A3B is a base pretrained checkpoint intended primarily for text completion and research. It may require careful prompting, fine-tuning, or post-processing for reliable assistant-style interactions. Moonlight-16B-A3B-Instruct is the more suitable variant for instruction following.


Sources 4
Provider

About Moonshot AI