Kimi Linear

Kimi-Linear-48B-A3B-Base

by Moonshot AI · Current open-weight model; publicly downloadable

Moonshot AI’s Kimi-Linear-48B-A3B-Base is an open-weight 48B-parameter mixture-of-experts foundation model with approximately 3B active parameters per token and a 1-million-token context window. Its Kimi Delta Attention and global Multi-Head Latent Attention architecture targets efficient long-context inference. The model is intended for local deployment, research, continued pretraining, and fine-tuning rather than turnkey chat or hosted API use.

Text Reasoning Coding
Kimi-Linear-48B-A3B-Base is the pretrained foundation checkpoint in Moonshot AI’s Kimi Linear family. It is designed for developers and researchers who need to process very large text inputs while retaining the ability to download, adapt, and run the model using compatible inference software. The checkpoint is open-weight and MIT licensed, but it is not instruction-tuned, so it should not be treated as a ready-made conversational assistant. Its main distinction is the combination of a million-token context window, sparse mixture-of-experts computation, and a hybrid attention architecture intended to make long-context inference more efficient.
Outputs

What Kimi-Linear-48B-A3B-Base can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Kimi Linear
Model type Lightweight
Context window 1.05M tokens
Release date 2025-10-30
Status Current open-weight model; publicly downloadable
Knowledge cutoff notes

No authoritative knowledge-cutoff date is stated in the exact model card or Kimi Linear technical report.

Model notes

This is the pretrained Base checkpoint rather than the instruction-tuned variant. It has 48B total parameters and approximately 3B activated parameters per token. Moonshot AI documents a 1M-token context window, Kimi Delta Attention combined with global Multi-Head Latent Attention, and deployment through Transformers, vLLM, and SGLang. The model card recommends Python 3.10+, PyTorch 2.6+, fla-core 0.4.0+, and trusted remote code. The Hugging Face card lists an MIT license. The reported up-to-75% KV-cache reduction and up-to-6x decoding improvement are architecture and benchmark claims that depend on context length, hardware, kernels, and serving configuration. No exact knowledge cutoff, maximum output-token limit, hosted pricing, official web-search integration, deprecation date, or shutdown date was found.

Cost

Model pricing

Input Not applicable; open-weight checkpoint with no official hosted API price
Output Not applicable; open-weight checkpoint with no official hosted API price
Model guide

Kimi-Linear-48B-A3B-Base: An Open-Weight Model for Million-Token Contexts

Kimi-Linear-48B-A3B-Base is Moonshot AI’s open-weight foundation model for long-context text generation. Its 48-billion-parameter mixture-of-experts design activates approximately 3 billion parameters per token and combines Kimi Delta Attention with global Multi-Head Latent Attention to reduce attention-cache demands during long-context inference. It supports up to 1,048,576 tokens of context and is intended for local deployment, research, continued pretraining, fine-tuning, and custom applications rather than direct consumer chat or a hosted API.

What Kimi-Linear-48B-A3B-Base is

Kimi-Linear-48B-A3B-Base is an open-weight text-generation model from Moonshot AI. The word “Base” identifies it as a pretrained foundation checkpoint rather than an instruction-tuned assistant. In practical terms, it is a model that developers can use as a starting point for continued pretraining, supervised fine-tuning, evaluation, and custom inference workflows.

The model contains approximately 48 billion total parameters, but it uses a mixture-of-experts design in which approximately 3 billion parameters are activated for each token. This can reduce the computation required for an individual token compared with a dense model of the same total size. It does not, however, mean that the model only requires memory for 3 billion parameters: deployment still needs to account for the full checkpoint, routing components, runtime overhead, and the selected numerical precision.

Moonshot AI released the model alongside the Kimi Linear technical report on October 30, 2025. The Hugging Face model card identifies the weights as MIT licensed and provides the primary distribution point.

Why the hybrid attention architecture matters

Kimi-Linear-48B-A3B-Base combines Kimi Delta Attention, abbreviated KDA, with global Multi-Head Latent Attention, or MLA. KDA is designed to maintain a compact recurrent-style state while processing a sequence, whereas the global MLA layers provide periodic broad information exchange across the context. The documented architecture uses a reported three-to-one ratio of KDA layers to global MLA layers.

This design targets a common problem in long-context language models: the key-value cache can become very large as the input grows. The key-value cache stores information needed to attend to earlier tokens during generation. Moonshot AI reports that Kimi Linear can reduce KV-cache requirements by up to 75% and achieve up to six times faster decoding in long-context settings. These are provider-reported architectural and benchmark claims, not guaranteed results. Actual memory use and speed depend on context length, hardware, quantization, kernels, batch size, and the inference engine.

The practical implication is that the model is especially interesting for workloads involving very long documents, large code repositories, extended agent traces, or other inputs where conventional full-attention inference becomes expensive. The architecture does not remove the need for substantial hardware, particularly when using the full-precision checkpoint.

Context window and output limits

The documented context length is 1,048,576 tokens, commonly described as a 1-million-token context window. This is the model’s supported maximum context length, not a promise that every deployment will run that size efficiently. Serving software, available memory, batching, prompt structure, and the chosen precision can impose lower practical limits.

No authoritative maximum output-token limit is specified for this exact checkpoint in the supplied model documentation. The total input-plus-output budget should therefore be checked against the configuration of the particular Transformers, vLLM, or SGLang deployment rather than assumed from the headline context figure. No authoritative knowledge-cutoff date is stated for the model card or technical report.

Capabilities and supported modalities

This is a text-in, text-out language model. It accepts text and produces text. It is not documented as a vision, audio, video, image-generation, speech, embedding, or music model. It also does not provide a first-party web-search tool or native external-action system.

  • Long-context text generation: Suitable for large documents, codebases, research collections, and extended text histories.
  • Continued pretraining: The base checkpoint can serve as a starting point for additional domain or language training.
  • Fine-tuning: Developers can adapt it for specialized generation or instruction-following behavior.
  • Code generation: It can generate and transform code as a general text model, but the supplied research does not identify a separate coding specialization or benchmark result.
  • Reasoning: It can perform reasoning through text generation, but it is not documented as a dedicated reasoning model with a separate reasoning mode or guaranteed chain-of-thought behavior.
  • Tool use: Native tool or function-calling support is not specified. Applications can connect the model to tools externally, but that should not be confused with built-in tool execution.

Because the checkpoint is not instruction-tuned, it may produce less predictable assistant-style answers than an instruction-tuned sibling. Developers who need direct chat behavior may need additional supervised training, a suitable prompt format, or a different model variant.

Deployment and hardware considerations

Moonshot AI provides the checkpoint through Hugging Face in Safetensors format. The documented Transformers workflow uses trusted remote code and recommends Python 3.10 or newer, PyTorch 2.6 or newer, and the fla-core package. The model documentation also describes deployment with vLLM and SGLang, including OpenAI-compatible local HTTP endpoints.

The checkpoint is described as a 48-billion-parameter BF16/F32 model. In practice, that generally places it in multi-GPU or quantized-deployment territory rather than ordinary consumer-laptop use. Quantization may reduce memory requirements, but the supplied research does not establish a particular quantized release, quality level, or minimum hardware configuration. The approximately 3-billion-parameter active count can reduce per-token computation, but it does not eliminate the need to store or access the full model structure.

Local serving can be useful when an organization needs control over data handling, custom fine-tuning, or integration with its own infrastructure. It also transfers responsibility for GPU capacity, software compatibility, monitoring, updates, and operational reliability to the deployer.

Pricing and API availability

There is no official hosted API price specified for Kimi-Linear-48B-A3B-Base. The checkpoint itself is downloadable as open weight, so the direct model price is not an input-token or output-token fee. Running it still creates infrastructure costs for GPUs, storage, electricity, engineering, and serving operations.

This distinction matters because the model is not the same as using a hosted Kimi consumer product or a commercial Kimi API endpoint. The supplied research does not identify a first-party hosted API offering, batch API, guaranteed structured-output mode, or provider-managed service for this exact Base checkpoint. Local deployments may expose streaming responses through compatible serving systems, but the behavior depends on the selected server and configuration.

Main strengths and trade-offs

The clearest strength is its positioning for very long text inputs. A million-token context can accommodate unusually large source material in a single request, and the hybrid attention design is intended to make long-context decoding less demanding than a conventional full-attention approach. The open-weight release also gives technical teams more control than a closed hosted model: weights can be inspected, adapted, evaluated, and deployed in a private environment.

The main trade-off is operational complexity. A 48-billion-parameter checkpoint is still a substantial model despite sparse activation. Specialized attention support, trusted remote code, current framework versions, and compatible kernels are part of the deployment burden. The model’s long-context capability may also be unnecessary for short prompts, where a smaller instruction-tuned model can be easier and cheaper to operate.

Moonshot AI’s reported cache and decoding improvements should be treated as claims to validate in the intended environment. They do not establish that every workload will be six times faster or that memory use will always fall by 75%.

When to choose Kimi-Linear-48B-A3B-Base

This model is a reasonable choice when the primary requirement is a downloadable foundation model capable of handling very long text contexts. It is particularly relevant for:

  • Research into efficient attention and mixture-of-experts architectures.
  • Processing large technical, legal, scientific, or organizational document collections.
  • Analysis of large codebases or long software-development histories.
  • Continued pretraining on a specialized corpus.
  • Fine-tuning a long-context model for a custom text-generation task.
  • Private or local inference where using a hosted service is unsuitable.

Another option may be more appropriate when the priority is immediate assistant-style interaction, built-in tools, multimodal input, hosted billing, or a smaller hardware footprint. The corresponding instruction-tuned Kimi Linear checkpoint is the more natural comparison for conversational behavior, while a smaller dense or specialized model may be preferable for short-context, low-cost, or latency-sensitive workloads. Those alternatives may sacrifice some of the Base model’s long-context or customization advantages, but they can reduce setup effort and improve out-of-the-box usability.

Limitations to check before deployment

Kimi-Linear-48B-A3B-Base is not an instruction-tuned chatbot, so raw prompting may not produce reliable assistant behavior. It has no documented native vision, audio, video, image-generation, or web-search capability. Structured-output guarantees, a maximum generation limit, a knowledge-cutoff date, hosted pricing, deprecation timing, and shutdown timing are not specified for this exact checkpoint.

Developers should also verify the license and model-card requirements in the intended commercial or research setting, test the exact inference engine with representative context lengths, and measure quality after any quantization or fine-tuning. The model’s headline context window is valuable only if the deployment can support it within acceptable memory, latency, and cost limits.

Bottom line

Kimi-Linear-48B-A3B-Base is best understood as a long-context research and deployment foundation model, not a finished chatbot or a conventional hosted API product. Its combination of 48 billion total parameters, approximately 3 billion active parameters per token, hybrid linear/global attention, and a 1-million-token context window makes it distinctive for teams willing to manage substantial local infrastructure. Its open-weight MIT-licensed release supports customization, but users seeking turnkey conversation, native tools, multimodal input, or simple pay-as-you-go access should consider a different model type.


Answers to Frequently Asked Questions

What hardware and software are needed to run Kimi-Linear-48B-A3B-Base?
Running the approximately 48-billion-parameter checkpoint generally requires a multi-GPU or quantized deployment rather than a typical consumer laptop. The documented setup uses Safetensors, Python 3.10 or newer, PyTorch 2.6 or newer, the fla-core package, and can support serving through Transformers, vLLM, or SGLang. Exact requirements depend on precision, context length, quantization, and serving configuration.
Is Kimi-Linear-48B-A3B-Base an instruction-tuned chatbot?
No. The “Base” designation means it is a pretrained foundation checkpoint rather than an instruction-tuned assistant. Developers may need additional supervised fine-tuning, prompt formatting, or an instruction-tuned model variant for reliable conversational behavior.
What makes Kimi-Linear’s hybrid attention architecture different?
Kimi-Linear combines Kimi Delta Attention (KDA), which maintains a compact recurrent-style state, with global Multi-Head Latent Attention (MLA) layers that enable broader information exchange. This design is intended to reduce key-value cache requirements and improve long-context decoding efficiency compared with conventional full attention.
What is Kimi-Linear-48B-A3B-Base?
Kimi-Linear-48B-A3B-Base is an open-weight, pretrained foundation language model from Moonshot AI. It has approximately 48 billion total parameters, with about 3 billion activated per token through a mixture-of-experts design, and is intended for continued pretraining, fine-tuning, evaluation, and custom inference.
How large is Kimi-Linear-48B-A3B-Base’s context window?
The model supports a documented maximum context length of 1,048,576 tokens, commonly described as a 1-million-token context window. Actual usable context may be lower depending on the inference engine, hardware, precision, batching, and available memory.


Sources 4
Provider

About Moonshot AI