What Kimi-Linear-48B-A3B-Base is
Kimi-Linear-48B-A3B-Base is an open-weight text-generation model from Moonshot AI. The word “Base” identifies it as a pretrained foundation checkpoint rather than an instruction-tuned assistant. In practical terms, it is a model that developers can use as a starting point for continued pretraining, supervised fine-tuning, evaluation, and custom inference workflows.
The model contains approximately 48 billion total parameters, but it uses a mixture-of-experts design in which approximately 3 billion parameters are activated for each token. This can reduce the computation required for an individual token compared with a dense model of the same total size. It does not, however, mean that the model only requires memory for 3 billion parameters: deployment still needs to account for the full checkpoint, routing components, runtime overhead, and the selected numerical precision.
Moonshot AI released the model alongside the Kimi Linear technical report on October 30, 2025. The Hugging Face model card identifies the weights as MIT licensed and provides the primary distribution point.
Why the hybrid attention architecture matters
Kimi-Linear-48B-A3B-Base combines Kimi Delta Attention, abbreviated KDA, with global Multi-Head Latent Attention, or MLA. KDA is designed to maintain a compact recurrent-style state while processing a sequence, whereas the global MLA layers provide periodic broad information exchange across the context. The documented architecture uses a reported three-to-one ratio of KDA layers to global MLA layers.
This design targets a common problem in long-context language models: the key-value cache can become very large as the input grows. The key-value cache stores information needed to attend to earlier tokens during generation. Moonshot AI reports that Kimi Linear can reduce KV-cache requirements by up to 75% and achieve up to six times faster decoding in long-context settings. These are provider-reported architectural and benchmark claims, not guaranteed results. Actual memory use and speed depend on context length, hardware, quantization, kernels, batch size, and the inference engine.
The practical implication is that the model is especially interesting for workloads involving very long documents, large code repositories, extended agent traces, or other inputs where conventional full-attention inference becomes expensive. The architecture does not remove the need for substantial hardware, particularly when using the full-precision checkpoint.
Context window and output limits
The documented context length is 1,048,576 tokens, commonly described as a 1-million-token context window. This is the model’s supported maximum context length, not a promise that every deployment will run that size efficiently. Serving software, available memory, batching, prompt structure, and the chosen precision can impose lower practical limits.
No authoritative maximum output-token limit is specified for this exact checkpoint in the supplied model documentation. The total input-plus-output budget should therefore be checked against the configuration of the particular Transformers, vLLM, or SGLang deployment rather than assumed from the headline context figure. No authoritative knowledge-cutoff date is stated for the model card or technical report.
Capabilities and supported modalities
This is a text-in, text-out language model. It accepts text and produces text. It is not documented as a vision, audio, video, image-generation, speech, embedding, or music model. It also does not provide a first-party web-search tool or native external-action system.
- Long-context text generation: Suitable for large documents, codebases, research collections, and extended text histories.
- Continued pretraining: The base checkpoint can serve as a starting point for additional domain or language training.
- Fine-tuning: Developers can adapt it for specialized generation or instruction-following behavior.
- Code generation: It can generate and transform code as a general text model, but the supplied research does not identify a separate coding specialization or benchmark result.
- Reasoning: It can perform reasoning through text generation, but it is not documented as a dedicated reasoning model with a separate reasoning mode or guaranteed chain-of-thought behavior.
- Tool use: Native tool or function-calling support is not specified. Applications can connect the model to tools externally, but that should not be confused with built-in tool execution.
Because the checkpoint is not instruction-tuned, it may produce less predictable assistant-style answers than an instruction-tuned sibling. Developers who need direct chat behavior may need additional supervised training, a suitable prompt format, or a different model variant.
Deployment and hardware considerations
Moonshot AI provides the checkpoint through Hugging Face in Safetensors format. The documented Transformers workflow uses trusted remote code and recommends Python 3.10 or newer, PyTorch 2.6 or newer, and the fla-core package. The model documentation also describes deployment with vLLM and SGLang, including OpenAI-compatible local HTTP endpoints.
The checkpoint is described as a 48-billion-parameter BF16/F32 model. In practice, that generally places it in multi-GPU or quantized-deployment territory rather than ordinary consumer-laptop use. Quantization may reduce memory requirements, but the supplied research does not establish a particular quantized release, quality level, or minimum hardware configuration. The approximately 3-billion-parameter active count can reduce per-token computation, but it does not eliminate the need to store or access the full model structure.
Local serving can be useful when an organization needs control over data handling, custom fine-tuning, or integration with its own infrastructure. It also transfers responsibility for GPU capacity, software compatibility, monitoring, updates, and operational reliability to the deployer.
Pricing and API availability
There is no official hosted API price specified for Kimi-Linear-48B-A3B-Base. The checkpoint itself is downloadable as open weight, so the direct model price is not an input-token or output-token fee. Running it still creates infrastructure costs for GPUs, storage, electricity, engineering, and serving operations.
This distinction matters because the model is not the same as using a hosted Kimi consumer product or a commercial Kimi API endpoint. The supplied research does not identify a first-party hosted API offering, batch API, guaranteed structured-output mode, or provider-managed service for this exact Base checkpoint. Local deployments may expose streaming responses through compatible serving systems, but the behavior depends on the selected server and configuration.
Main strengths and trade-offs
The clearest strength is its positioning for very long text inputs. A million-token context can accommodate unusually large source material in a single request, and the hybrid attention design is intended to make long-context decoding less demanding than a conventional full-attention approach. The open-weight release also gives technical teams more control than a closed hosted model: weights can be inspected, adapted, evaluated, and deployed in a private environment.
The main trade-off is operational complexity. A 48-billion-parameter checkpoint is still a substantial model despite sparse activation. Specialized attention support, trusted remote code, current framework versions, and compatible kernels are part of the deployment burden. The model’s long-context capability may also be unnecessary for short prompts, where a smaller instruction-tuned model can be easier and cheaper to operate.
Moonshot AI’s reported cache and decoding improvements should be treated as claims to validate in the intended environment. They do not establish that every workload will be six times faster or that memory use will always fall by 75%.
When to choose Kimi-Linear-48B-A3B-Base
This model is a reasonable choice when the primary requirement is a downloadable foundation model capable of handling very long text contexts. It is particularly relevant for:
- Research into efficient attention and mixture-of-experts architectures.
- Processing large technical, legal, scientific, or organizational document collections.
- Analysis of large codebases or long software-development histories.
- Continued pretraining on a specialized corpus.
- Fine-tuning a long-context model for a custom text-generation task.
- Private or local inference where using a hosted service is unsuitable.
Another option may be more appropriate when the priority is immediate assistant-style interaction, built-in tools, multimodal input, hosted billing, or a smaller hardware footprint. The corresponding instruction-tuned Kimi Linear checkpoint is the more natural comparison for conversational behavior, while a smaller dense or specialized model may be preferable for short-context, low-cost, or latency-sensitive workloads. Those alternatives may sacrifice some of the Base model’s long-context or customization advantages, but they can reduce setup effort and improve out-of-the-box usability.
Limitations to check before deployment
Kimi-Linear-48B-A3B-Base is not an instruction-tuned chatbot, so raw prompting may not produce reliable assistant behavior. It has no documented native vision, audio, video, image-generation, or web-search capability. Structured-output guarantees, a maximum generation limit, a knowledge-cutoff date, hosted pricing, deprecation timing, and shutdown timing are not specified for this exact checkpoint.
Developers should also verify the license and model-card requirements in the intended commercial or research setting, test the exact inference engine with representative context lengths, and measure quality after any quantization or fine-tuning. The model’s headline context window is valuable only if the deployment can support it within acceptable memory, latency, and cost limits.
Bottom line
Kimi-Linear-48B-A3B-Base is best understood as a long-context research and deployment foundation model, not a finished chatbot or a conventional hosted API product. Its combination of 48 billion total parameters, approximately 3 billion active parameters per token, hybrid linear/global attention, and a 1-million-token context window makes it distinctive for teams willing to manage substantial local infrastructure. Its open-weight MIT-licensed release supports customization, but users seeking turnkey conversation, native tools, multimodal input, or simple pay-as-you-go access should consider a different model type.

