Moonlight

Moonlight-16B-A3B-Instruct

by Moonshot AI · Available open-weight checkpoint

Moonlight-16B-A3B-Instruct is Moonshot AI’s open-weight, MIT-licensed instruction model for self-hosted text generation. Its MoE architecture contains approximately 16B total parameters while activating about 3B per token. It supports an 8,192-token context window, text input and output, and deployment with Transformers or vLLM. No first-party hosted pricing or verified native multimodal, tool-use, structured-output, or separate maximum-output capabilities were identified.

Text Reasoning Coding
Moonlight-16B-A3B-Instruct is Moonshot AI’s open instruction-tuned language model for users who want downloadable weights rather than a managed chat subscription or first-party hosted API. It combines a 16-billion-parameter mixture-of-experts design with approximately 3 billion active parameters per token, reducing the computation used for each generated token compared with a similarly sized dense model. The checkpoint is intended for local and self-hosted text generation, with an 8,192-token context window and an MIT license.
Outputs

What Moonlight-16B-A3B-Instruct can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

6/10 Reasoning
6/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Moonlight
Model type General Purpose
Context window 8K tokens
Release date 2025-02-24
Status Available open-weight checkpoint
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published in the reviewed model card, technical report, or repository documentation.

Model notes

Open-weight instruction-tuned checkpoint hosted as moonshotai/Moonlight-16B-A3B-Instruct under the MIT license. The model has approximately 16B total parameters and approximately 3B activated parameters per token. Moonshot AI describes the Moonlight family as a mixture-of-experts model trained on 5.7T tokens using an improved Muon optimizer. The model card identifies an 8K context length and recommends Transformers with trust_remote_code=True; vLLM deployment is also documented. No official first-party hosted API price, separate maximum-output limit, knowledge cutoff, web-search integration, or exact model-level support for structured output, JSON mode, caching, batching, streaming, or fine-tuning was verified. Editorial scores are comparative estimates rather than vendor-provided ratings.

Model guide

Moonlight-16B-A3B-Instruct: Open 3B-Active MoE for Self-Hosted Text Generation

Moonlight-16B-A3B-Instruct is an open-weight, MIT-licensed instruction-tuned language model from Moonshot AI. Its mixture-of-experts architecture contains approximately 16 billion total parameters but activates about 3 billion per token, making it a practical candidate for local or self-hosted text generation. It has an 8,192-token context window and supports text input and text output, but no verified native image, audio, video, tool-use, hosted API pricing, or separate maximum-output specification is documented for this checkpoint.

What is Moonlight-16B-A3B-Instruct?

Moonlight-16B-A3B-Instruct is an instruction-tuned, open-weight language model released by Moonshot AI. Instruction tuning means the model has been trained to respond to user directions rather than only predict text from a continuation prompt. In practical terms, it is intended for tasks such as drafting, question answering, summarization, general text generation, coding assistance, and research experimentation.

The checkpoint was released on February 24, 2025, and is hosted on Hugging Face as moonshotai/Moonlight-16B-A3B-Instruct. Its weights are available under the MIT license, subject to the terms and conditions of that license. This makes it different from a consumer-only model accessed through a web application: operators can download the checkpoint and manage inference themselves, provided they supply suitable hardware and an appropriate serving stack.

Moonlight-16B-A3B-Instruct belongs to Moonshot AI’s Moonlight model family. The related moonshotai/Moonlight-16B-A3B checkpoint is a pretrained version, whereas the model covered here is the instruction-tuned variant intended for direct interaction and instruction following.

Architecture and compute profile

The model uses a mixture-of-experts, or MoE, architecture. Rather than applying all of its parameters to every token, an MoE model routes each token to a smaller set of specialized expert components. Moonshot AI describes Moonlight as a 3B/16B MoE model: it has approximately 16 billion parameters in total, while approximately 3 billion are activated for each token.

The distinction matters for deployment. The total parameter count affects memory requirements because the model’s weights still need to be stored, while the active-parameter count is more closely related to the computation required for each token. As a result, Moonlight-16B-A3B-Instruct may offer a useful middle ground between a small dense model and a much larger dense model. It is not equivalent to a 3-billion-parameter model in every respect, however; operators still need to account for the full checkpoint, runtime overhead, precision, and any quantization used during deployment.

According to Moonshot AI’s description of the Moonlight family, the model was trained on 5.7 trillion tokens using an improved Muon optimizer. The architecture is based on the DeepSeek-V3 design. These are provider or project claims about the model’s development and should not be interpreted as a guarantee of performance for a particular workload.

Context window and supported modalities

The verified context length is 8,192 tokens. A token is a piece of text processed by the model; the context window includes the prompt and the relevant generated continuation within the model’s supported sequence length. An 8K window is sufficient for ordinary questions, short documents, code snippets, and compact research notes, but it is relatively limited for workflows involving long books, large repositories, or many files in one prompt.

Moonlight-16B-A3B-Instruct is a text model. It accepts text input and produces text output. The reviewed model documentation does not identify native image, audio, or video input, and it does not document image, audio, video, speech, music, embedding, or other non-text output. Images or files cannot therefore be assumed to work simply because a surrounding application might offer multimodal features.

CapabilityVerified status
Text inputSupported
Text outputSupported
Context length8,192 tokens
Image, audio, or video inputNot documented for this checkpoint
Image, audio, or video outputNot documented for this checkpoint
Knowledge cutoffNo authoritative date identified

Reasoning, coding, and general performance

Moonlight-16B-A3B-Instruct is positioned as a general-purpose instruction model rather than as a dedicated vision, speech, or tool-using system. The supplied research reports strong results for the model’s training scale across coding, mathematics, English, and Chinese benchmarks, but it does not provide the individual benchmark scores in the reviewed material. Those claims are best treated as reported model-card or project results rather than as a guarantee that the model will outperform every similarly sized alternative.

For coding, the model can be evaluated for code completion, explanation, transformation, and debugging prompts because it generates text and the project reports coding capability among its areas of strength. The available information does not verify repository-scale code understanding, an integrated execution environment, automatic testing, or a tool loop. Any code execution, file manipulation, web search, or external action must be supplied by the application around the model rather than assumed to be a native model feature.

The same distinction applies to reasoning. The model can produce step-by-step-looking text for mathematical or analytical prompts, and the project reports mathematics performance, but no separate reasoning mode or verified reasoning-token feature is documented. Users should assess factual reliability and problem-solving quality on their own tasks instead of treating the model’s MoE design or benchmark claims as proof of universal reasoning strength.

Deployment and availability

The canonical instruction-tuned checkpoint is available from Hugging Face under moonshotai/Moonlight-16B-A3B-Instruct. The model card documents loading it with Hugging Face Transformers and recommends using the repository’s custom model code with trust_remote_code=True. It can also be served with vLLM through an OpenAI-compatible local endpoint.

Because this is an open-weight checkpoint, deployment responsibilities move to the operator. Hardware capacity, model precision, quantization, batching, concurrency, latency targets, monitoring, security, and operating costs are not provided as a managed service by the checkpoint itself. The practical memory requirement will depend on the weight format and serving configuration, but the supplied research does not establish a single hardware minimum and one should not be inferred.

The OpenAI-compatible endpoint available through a local vLLM deployment describes an interface style, not official OpenAI hosting or a Moonshot AI hosted API. No first-party hosted API pricing was verified for this exact model. There is also no separately documented maximum-output-token limit. The 8,192-token context length should not automatically be interpreted as a guaranteed output length, since the context includes the prompt and the model documentation does not publish a distinct maximum completion value.

Pricing and operational costs

There is no verified per-token input or output price for a Moonshot AI hosted endpoint serving Moonlight-16B-A3B-Instruct. The checkpoint itself is open-weight and MIT-licensed, so there is no model-download fee identified in the supplied research. That does not mean deployment is free: users must pay for hardware, cloud instances, storage, electricity, bandwidth, maintenance, and engineering work where applicable.

This cost model can be attractive for predictable or privacy-sensitive workloads that justify operating an inference service. It can be less attractive for occasional users who would rather pay for an on-demand managed API and avoid maintaining infrastructure. The model’s approximately 3-billion active-parameter profile may improve compute efficiency relative to a dense model of comparable total size, but actual cost and speed depend on the serving engine, hardware, quantization, request length, and concurrency.

Main strengths and limitations

Its clearest strength is the combination of open access and a relatively efficient MoE compute profile. The MIT license and downloadable weights support experimentation, private deployments, customization of the surrounding application, and local inference without depending on a consumer chat interface. The model also provides a single text-focused checkpoint for English and Chinese work, general instruction following, coding experiments, and mathematical prompts.

  • Open deployment: Users can download and serve the checkpoint rather than relying on a closed hosted endpoint.
  • MoE efficiency: Approximately 3 billion parameters are activated per token even though the model contains approximately 16 billion total parameters.
  • Broad text use: The model is intended for general instruction following, text generation, research, mathematics, and coding-related tasks.
  • Permissive licensing: The checkpoint is distributed under the MIT license.
  • Self-hosting flexibility: Transformers, vLLM, and compatible inference systems are documented deployment routes.

The main limitations are equally important. The 8,192-token context is modest for long-document processing. Native multimodal input and output are not documented. There is no verified first-party hosted price, separate maximum-output specification, knowledge-cutoff date, or model-level guarantee for structured output, JSON mode, streaming, caching, batching, fine-tuning, or tool use. Some of these functions may be implemented by an inference engine or application wrapper, but that should not be confused with a verified capability of the checkpoint.

When to choose Moonlight-16B-A3B-Instruct

Choose Moonlight-16B-A3B-Instruct when you want an open, downloadable instruction model for local or self-hosted text generation and can operate the required infrastructure. It is a reasonable candidate for private prototypes, general assistants without native media requirements, Chinese or English text workflows, code-generation experiments, mathematical prompting, and applications where avoiding a managed per-token API is valuable.

Its MoE structure is especially relevant when you want more total model capacity than a small dense model while keeping per-token computation closer to a lower active-parameter class. The trade-off is that the complete model still has to be stored and served, so active parameters alone do not determine the hardware requirement.

Another option may be more appropriate if your application needs a context window substantially longer than 8K tokens, native image or audio understanding, image or video generation, built-in web search, a verified function-calling contract, or a managed service with published pricing and service-level expectations. A dedicated hosted model may also be preferable for teams that do not want to manage inference infrastructure. Conversely, a smaller dense model may be easier to run when low memory usage and simple deployment matter more than the capacity associated with a 16-billion-parameter MoE checkpoint.

Bottom line

Moonlight-16B-A3B-Instruct is best understood as an open, text-only, instruction-tuned MoE checkpoint for operators who want control over deployment. Its approximately 16-billion-parameter total size, approximately 3-billion active-parameter profile, MIT license, and documented Transformers and vLLM paths make it more suitable for self-hosted experimentation than for readers looking for a ready-made consumer service. The 8,192-token context window and unverified status of hosted pricing, tool use, structured output, and multimodal features define the boundaries of what can confidently be expected from this specific model.


Answers to Frequently Asked Questions

How can Moonlight-16B-A3B-Instruct be deployed?
The model can be loaded with Hugging Face Transformers using the repository’s custom model code and can also be served with vLLM through an OpenAI-compatible local endpoint. Operators must provide the required hardware and manage memory, precision, quantization, concurrency, monitoring, security, and operating costs.
Can Moonlight-16B-A3B-Instruct process images, audio, or video?
Moonlight-16B-A3B-Instruct is documented as a text-input and text-output model. Native image, audio, or video input and output are not documented for this checkpoint, so multimodal capabilities should not be assumed.
What is the context window of Moonlight-16B-A3B-Instruct?
The verified context window is 8,192 tokens, including the prompt and generated continuation. This is suitable for ordinary questions, short documents, code snippets, and compact research notes, but it may be limiting for long books, large repositories, or multi-file prompts.
How many parameters does Moonlight-16B-A3B-Instruct activate per token?
Moonlight-16B-A3B-Instruct is a mixture-of-experts model with approximately 16 billion total parameters and approximately 3 billion active parameters per token. The active-parameter count affects computation, while the full parameter count still affects memory requirements because the complete checkpoint must be stored.
What is Moonlight-16B-A3B-Instruct?
Moonlight-16B-A3B-Instruct is an open-weight, instruction-tuned language model from Moonshot AI designed for self-hosted text generation, question answering, summarization, coding assistance, mathematics, and research experimentation. Its weights are available on Hugging Face under the MIT license.


Sources 5
Provider

About Moonshot AI