Kimi Linear

Kimi-Linear-48B-A3B-Instruct

by Moonshot AI · Current open-weight instruction-tuned checkpoint

Moonshot AI's Kimi-Linear-48B-A3B-Instruct is an MIT-licensed, open-weight text model for long-context generation and self-hosted inference. It contains 48 billion total parameters, activates approximately 3 billion per token, supports up to 1,048,576 tokens of context, and uses a hybrid Kimi Delta Attention and Multi-Head Latent Attention architecture intended to reduce long-context memory use and improve decoding efficiency.

Text Reasoning Coding
Kimi-Linear-48B-A3B-Instruct is built for long-context text generation, conversation, document analysis, coding, and self-hosted inference. Its hybrid attention design is intended to reduce key-value cache memory requirements and improve decoding efficiency at very long context lengths, while its open-weight MIT-licensed distribution gives developers more control than a typical hosted-only model.
Outputs

What Kimi-Linear-48B-A3B-Instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Kimi Linear
Model type General Purpose
Context window 1.05M tokens
Release date 2025-10-30
Status Current open-weight instruction-tuned checkpoint
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was identified in the primary model card or Kimi Linear technical report.

Model notes

Open-weight MIT-licensed checkpoint hosted by Moonshot AI on Hugging Face. The model has 48B total parameters and approximately 3B activated parameters per token. Moonshot AI describes a hybrid architecture combining Kimi Delta Attention with global Multi-Head Latent Attention, with up to 75% lower KV-cache usage and up to 6x faster decoding in reported long-context comparisons. The 1M-token context length is stated in the model card. The exact model card does not publish a separate fixed maximum output-token limit, official hosted API pricing, knowledge cutoff, or a distinct legacy JSON-mode capability. Deployment examples use Transformers, vLLM, and SGLang; trust_remote_code may be required.

Model guide

Kimi-Linear-48B-A3B-Instruct: A 1M-Token Open-Weight Model for Efficient Long-Context Inference

Kimi-Linear-48B-A3B-Instruct is Moonshot AI's instruction-tuned open-weight language model. It combines Kimi Delta Attention with global Multi-Head Latent Attention, contains 48 billion total parameters, activates approximately 3 billion parameters per token, and supports a context window of up to 1,048,576 tokens.

What is Kimi-Linear-48B-A3B-Instruct?

Kimi-Linear-48B-A3B-Instruct is an instruction-tuned, open-weight language model from Moonshot AI. It is part of the Kimi Linear model family and is available as a checkpoint through Moonshot AI's Hugging Face repository under the MIT license. The instruction-tuned designation means that the model has been adapted to follow user requests and produce conversational or task-oriented responses rather than simply continuing arbitrary text.

The model is designed primarily for text input and text generation. Its most important distinction is the combination of a very large context window with a hybrid attention architecture intended to make long-context decoding more efficient. This makes it relevant to developers working with long documents, large code repositories, extended conversations, and research workloads that may exceed the practical context capacity of many smaller models.

Specifications and position in the Kimi lineup

Kimi-Linear-48B-A3B-Instruct has 48 billion total parameters, but Moonshot AI reports that approximately 3 billion parameters are activated for each token. In practical terms, the model retains the representational capacity of a much larger model while using a sparse activation pattern during inference. The total parameter count still matters for storage and deployment, so the active-parameter figure should not be interpreted as meaning that the checkpoint only requires the resources of a 3-billion-parameter model.

SpecificationVerified detail
ProviderMoonshot AI
Model familyKimi Linear
Model typeInstruction-tuned general-purpose language model
Total parameters48 billion
Active parametersApproximately 3 billion per token
Context lengthUp to 1,048,576 tokens
Input and outputText input and text generation output
LicenseMIT
Hosted model-specific pricingNot identified for this checkpoint

Within Moonshot AI's broader catalog, this checkpoint is a self-hosting-oriented Kimi model rather than a consumer chat feature or a clearly documented, model-specific hosted API product. The supplied research does not identify a separate Moonshot-hosted endpoint with published input and output prices for this exact checkpoint.

How the architecture supports long contexts

The model combines Kimi Delta Attention, described by Moonshot AI as a refined gated delta-rule linear-attention mechanism, with global Multi-Head Latent Attention. Conventional full attention compares tokens across a growing sequence, which can make memory use and decoding cost increasingly difficult to manage as context expands. Linear-attention components are designed to process sequence information with a more favorable memory pattern, while the global attention layers preserve broader information exchange across the sequence.

Kimi Linear uses a hybrid arrangement with many Kimi Delta Attention layers and a smaller number of global attention layers. This is intended to balance efficient sequence processing with the ability to connect information across distant parts of a document or conversation.

Moonshot AI reports up to 75% lower key-value cache requirements and up to six-times faster decoding at long context lengths compared with relevant full-attention baselines. These are provider-reported research results, not guarantees for every hardware setup, quantization method, sequence length, or serving framework. Actual performance will depend on GPU memory, parallelism, batch size, software versions, and the workload.

Context, input, and output limits

The model card states a maximum context length of up to 1,048,576 tokens, commonly described as a 1-million-token context window. The context window covers the material supplied to the model together with the generated response, so a very long prompt leaves less room for output within the same request.

No separate fixed maximum output-token limit was identified in the supplied primary documentation. That value should therefore be treated as unknown rather than assumed to equal the full context size. Serving software may impose its own response-length or memory limits.

A 1-million-token context window does not mean that every prompt will be equally useful at that length. Very large inputs can increase processing time, memory pressure, and operational cost. Developers should still test retrieval, document chunking, prompt organization, and output limits for their particular application.

Capabilities and supported modalities

Kimi-Linear-48B-A3B-Instruct is a text-only model. It accepts text and generates text; the supplied research does not support native image, audio, or video input or output. It is therefore not the appropriate checkpoint for multimodal understanding, image creation, speech generation, or video production.

  • Text generation: Supported.
  • Instruction following and conversation: Supported as primary uses of the instruction-tuned checkpoint.
  • Long-document processing: Supported, with a stated context capacity of up to 1 million tokens.
  • Coding: Suitable for code and repository-scale text workflows, although no model-specific coding benchmark is supplied.
  • Reasoning: Useful for multi-step language tasks, but no authoritative model-specific reasoning benchmark is provided.
  • Tool or function calling: Native tool-use support is not identified in the supplied research.
  • Structured or JSON mode: A distinct legacy JSON-mode capability is not documented.
  • Streaming: Supported through compatible inference-server deployments according to the supplied model data.

Coding and reasoning should be understood as practical use cases, not as provider-certified scores. The model can generate code and work through textual problems, but the supplied sources do not establish a particular ranking against other coding or reasoning models.

Deployment options and hardware considerations

The checkpoint can be used with Transformers and is also documented for serving through vLLM and SGLang. These serving systems can expose OpenAI-compatible HTTP endpoints, which may simplify integration with applications that already use a chat-completions-style client. The model card's deployment instructions require a recent PyTorch environment, Python 3.10 or newer, the Kimi Linear implementation, and the recommended fla-core package. Depending on the integration, remote model code may need to be trusted.

Despite activating only about 3 billion parameters per token, the model's 48-billion-parameter checkpoint remains substantial. Loading the weights, maintaining the context state, and serving multiple users can require significant GPU memory. Quantization may reduce memory requirements, while tensor parallelism can distribute the model across multiple accelerators. Both choices can affect compatibility, throughput, numerical behavior, and response quality, so they should be validated with the intended serving stack.

The architecture's efficiency is most relevant at long context lengths. For short prompts, the advantages over a smaller dense model may be less pronounced, while the operational burden of deploying a 48-billion-parameter checkpoint remains. This creates an important trade-off: the model may be attractive when context capacity and local control matter more than the simplest or cheapest deployment.

Pricing and access

No official Moonshot-hosted API price was identified for Kimi-Linear-48B-A3B-Instruct in the supplied research. It is distributed as an open-weight checkpoint, so there is no verified per-token price to quote for the model itself. Self-hosting costs instead come from hardware, electricity, storage, orchestration, maintenance, and any infrastructure used to serve it.

That does not necessarily make every deployment inexpensive. A large model may be economical for an organization with existing accelerator capacity and a high volume of long-context requests, but less economical for occasional users who would need to rent substantial GPU capacity. Cost comparisons should include utilization and memory requirements rather than focusing only on the absence of an API fee.

Main strengths and limitations

Strengths

  • Very large context capacity: The 1,048,576-token context window is useful for long documents, extensive codebases, and persistent textual workflows.
  • Long-context efficiency focus: The hybrid attention design targets lower key-value cache use and faster decoding as sequences become very long.
  • Open deployment: The MIT license and available checkpoint support local, private, research, and customized deployments, subject to the implementer's review of the license and operating environment.
  • Sparse activation: Approximately 3 billion parameters are activated per token despite 48 billion total parameters.
  • Serving ecosystem: Transformers, vLLM, and SGLang provide several routes for experimentation and production deployment.

Limitations

  • Large deployment footprint: The total checkpoint size and long-context state can still require substantial accelerator memory.
  • Text-only operation: It cannot natively replace a vision-language, image-generation, speech, or video model.
  • Unknown output ceiling: A separate fixed maximum output-token limit is not documented in the supplied sources.
  • No verified hosted price: Users looking for a turnkey endpoint with predictable per-token billing will need another option or their own serving arrangement.
  • Unverified advanced interfaces: Native tool calling and a distinct JSON mode are not established by the supplied research.
  • Performance depends on deployment: Provider-reported efficiency gains may not appear identically across different GPUs, quantization settings, context lengths, or frameworks.

When to choose this model

Choose Kimi-Linear-48B-A3B-Instruct when the application benefits from a very large textual context and the team can operate an open-weight model. Good examples include analyzing a collection of lengthy reports, exploring a large software repository, building a private document assistant, researching efficient attention mechanisms, or running long-lived text workflows without sending prompts to a proprietary hosted service.

It is especially compelling when long-context memory usage is a central engineering concern. The reported attention design may offer a better efficiency profile than a comparable full-attention model at large sequence lengths, although that claim should be tested on the target hardware.

Another model type may be more appropriate when the priority is a small, inexpensive deployment for short prompts, a fully managed API with published pricing, reliable native function calling, or multimodal input and output. A conventional smaller model may also be preferable when the application rarely uses long contexts and does not justify the memory and operational complexity of a 48-billion-parameter checkpoint.

Bottom line

Kimi-Linear-48B-A3B-Instruct is a specialized open-weight choice for developers who value long-context text processing, local control, and an architecture designed for more efficient decoding at scale. Its 48-billion-parameter size, approximately 3-billion active-parameter pattern, MIT license, and stated 1-million-token context window make it substantially different from a small general-purpose local model. The trade-off is deployment complexity: hardware requirements remain significant, hosted pricing is not established for this checkpoint, and multimodal, tool-use, and fixed output-limit capabilities are not documented as part of the supplied specification.


Answers to Frequently Asked Questions

Does Kimi-Linear-48B-A3B-Instruct support images, audio, or tool calling?
No native multimodal capability is documented for this checkpoint. It is a text-only model, and the supplied research does not establish native tool calling or a distinct JSON mode. It is primarily intended for text generation, conversation, coding, and long-document processing.
Can Kimi-Linear-48B-A3B-Instruct be self-hosted?
Yes. The model is available as an open-weight checkpoint under the MIT license and can be used with Transformers, vLLM, and SGLang. However, its 48-billion-parameter size and long-context state can require substantial GPU memory, so quantization or tensor parallelism may be necessary.
How large is Kimi-Linear-48B-A3B-Instruct's context window?
The model supports a maximum context length of up to 1,048,576 tokens, commonly described as a 1-million-token context window. This limit includes both the input prompt and the generated response.
What makes Kimi-Linear-48B-A3B-Instruct efficient for long-context inference?
The model combines Kimi Delta Attention, a linear-attention mechanism, with global Multi-Head Latent Attention. Moonshot AI reports up to 75% lower key-value cache requirements and up to six-times faster decoding at long context lengths compared with relevant full-attention baselines, although actual results depend on hardware and deployment settings.
What is Kimi-Linear-48B-A3B-Instruct?
Kimi-Linear-48B-A3B-Instruct is an instruction-tuned, open-weight language model from Moonshot AI. It has 48 billion total parameters, activates approximately 3 billion parameters per token, supports text input and generation, and is designed for efficient long-context inference.


Sources 4
Provider

About Moonshot AI