Kimi K2

Kimi K2 Base

by Moonshot AI · Open-weight and downloadable; legacy relative to newer Kimi K2.x models but still accessible from the official Hugging Face repository

Kimi K2 Base is Moonshot AI’s pretrained 1T-parameter mixture-of-experts foundation model for fine-tuning, research, and self-hosted language systems. It provides a 128K-token context window and modified MIT licensing, but its approximately 1.03 TB checkpoint, high hardware requirements, lack of verified hosted pricing, and absence of documented native multimodal features make it unsuitable for lightweight or turnkey deployments.

Text Reasoning Coding
Kimi K2 Base is the pretrained foundation variant of Moonshot AI’s Kimi K2 model family. It combines 1 trillion total parameters with approximately 32 billion activated parameters per token, a 128K-token context window, and a modified MIT license. Unlike Kimi K2 Instruct, it is intended primarily as a starting point for fine-tuning, research, and custom deployment rather than as a ready-made conversational assistant.
Outputs

What Kimi K2 Base can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
3/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Kimi K2
Model type General Purpose
Context window 131K tokens
Release date 2025-07-11
Status Open-weight and downloadable; legacy relative to newer Kimi K2.x models but still accessible from the official Hugging Face repository
Knowledge cutoff notes

Moonshot AI does not publish a directly verified knowledge-cutoff date for the Kimi K2 Base checkpoint in the official model card.

Model notes

Kimi K2 Base is the pretrained foundation checkpoint, distinct from Kimi K2 Instruct and Kimi K2 Thinking. Moonshot AI documents it as a starting point for fine-tuning and custom solutions. The architecture has 1T total parameters, approximately 32B activated parameters, 384 experts, top-8 expert selection, 61 layers, and a 128K context window. The official Hugging Face repository is approximately 1.03 TB and uses a modified MIT license. Official deployment examples cover Transformers, vLLM, SGLang, KTransformers, and TensorRT-LLM-compatible environments. Tool use, structured outputs, JSON mode, caching, batch API support, maximum output tokens, and knowledge cutoff are not separately verified for this exact Base checkpoint. Editorial scores are comparative estimates, not vendor-published ratings.

Cost

Model pricing

Input No official hosted API price for the Base checkpoint; self-hosted weights
Output No official hosted API price for the Base checkpoint; self-hosted weights
Model guide

Kimi K2 Base: An Open-Weight Foundation Model for Fine-Tuning

Kimi K2 Base is Moonshot AI’s pretrained, open-weight foundation model for researchers and developers who want to fine-tune or self-host a large language model. Its trillion-parameter mixture-of-experts design, 128K-token context window, and modified MIT license provide substantial flexibility, but its roughly 1.03 TB checkpoint and demanding infrastructure requirements make it unsuitable for small-scale or low-latency deployments.

What is Kimi K2 Base?

Kimi K2 Base is Moonshot AI’s pretrained foundation checkpoint in the Kimi K2 family. It is distributed as downloadable open weights through Hugging Face, giving research teams and infrastructure developers control over model serving, fine-tuning, prompting, and application-level safeguards.

The word “Base” is important. This model is not presented as a finished chat assistant. A base model has learned general language patterns and capabilities during pretraining, but it may not consistently follow conversational instructions, format responses for end users, or perform reliable tool calls without additional post-training and application logic. Moonshot AI positions Kimi K2 Base as a starting point for custom systems, while Kimi K2 Instruct is the more appropriate sibling for general-purpose chat and agentic interactions.

The model was released on July 11, 2025. The supplied model information describes it as open-weight and downloadable, although relatively legacy compared with newer Kimi K2.x models. It remains available through its official model repository.

Architecture and 128K-token context window

Kimi K2 Base uses a sparse mixture-of-experts architecture, commonly abbreviated as MoE. Instead of activating every parameter for every token, the model routes each token to a selected group of specialist neural networks. This allows a very large total parameter count without activating the entire model on every step, although the complete checkpoint still requires substantial storage and serving infrastructure.

  • Total parameters: 1 trillion
  • Activated parameters: approximately 32 billion per token
  • Layers: 61
  • Experts: 384, with eight selected experts per token and one shared expert
  • Attention heads: 64
  • Vocabulary: 160,000 tokens
  • Context length: 131,072 tokens, commonly described as 128K
  • Checkpoint format: FP8

A 128K-token context window allows a deployment to provide much more text to the model in one request than a conventional short-context system. Practical examples include large code repositories, lengthy technical documentation, research collections, or multi-document analysis. The context limit is not the same as a guaranteed output length: the supplied specifications do not verify a separate maximum output-token limit for this checkpoint.

The model card also identifies Multi-head Latent Attention and SwiGLU activation among the architecture components. The official repository is approximately 1.03 TB, so the context window should not be mistaken for an indication that the model is easy to run locally. Storage, memory, networking, quantization choices, and multi-GPU serving capacity all matter when deploying it.

Capabilities and supported modalities

Kimi K2 Base is a text-generation model. The supplied specifications verify text input and text output, but do not document native image, audio, or video input or output for this exact Base checkpoint. It should therefore be evaluated as a language-only foundation model rather than as a multimodal model.

Moonshot AI’s wider Kimi K2 materials associate the model family with knowledge, reasoning, coding, and agentic workloads. For Kimi K2 Base specifically, the research supports treating these as foundation-model capabilities rather than as a complete end-user agent package. The Base checkpoint can provide a strong starting point for coding systems, domain adaptation, and experimental reasoning workflows, but developers should add instruction tuning, tool orchestration, validation, and safety controls where those functions are required.

The comparative editorial assessment supplied for this model rates reasoning at 8 out of 10 and coding at 9 out of 10. These are editorial scores, not vendor-published benchmark results. They indicate the model’s perceived suitability for technical and programming work, but they should not be used as a substitute for testing on a team’s own tasks.

Fine-tuning and custom deployment

The primary reason to choose Kimi K2 Base is control. Teams can use the checkpoint as the foundation for a specialized language model, adapting it to a domain, response style, internal terminology, or task format. Possible projects include domain-specific assistants, coding experiments, research systems, and self-hosted applications where the operator needs to manage the model rather than rely on a finished hosted chatbot.

Official deployment material references Transformers, vLLM, SGLang, KTransformers, and TensorRT-LLM-compatible infrastructure. The model uses custom model code, so some serving approaches may require model-specific integration or remote-code settings. Operators should review the official repository and their selected serving framework before committing to an infrastructure design.

Self-hosting also shifts responsibility to the deploying organization. The model does not automatically provide a complete safety policy, reliable instruction-following layer, access control, monitoring system, or application-specific evaluation process. Those elements must be built around the checkpoint. The large repository size and high hardware requirements make this a poor fit for a small workstation or a deployment where fast responses are more important than model control.

Tool use, structured output, and API availability

There is no separately verified native tool-use or function-calling capability for Kimi K2 Base in the supplied research. The same is true for JSON mode, structured outputs, caching, batch API access, and a provider-hosted API price for this exact checkpoint. A developer could build an orchestration layer around a self-hosted model, but that would be an application decision rather than a confirmed built-in feature of the Base model.

The model card includes OpenAI-compatible serving examples. In context, these examples describe locally hosted inference and should not be interpreted as proof of a guaranteed Moonshot-hosted API product for Kimi K2 Base. Organizations considering the model should distinguish between running an OpenAI-compatible server themselves and purchasing a managed endpoint from the model provider.

Pricing and operating cost

Kimi K2 Base has no verified official hosted API price in the supplied information. It is distributed as self-hosted weights, so the direct model price is not presented as a recurring per-token subscription or standard hosted inference rate.

That does not make deployment free. The effective cost includes storage for a roughly 1.03 TB repository, accelerator hardware or rented GPU capacity, electricity, networking, engineering time, serving software, monitoring, and ongoing maintenance. The editorial cost score is 6 out of 10, while the speed score is 3 out of 10; both are comparative estimates rather than Moonshot AI specifications. In practical terms, this model’s sparse architecture may reduce per-token activation relative to a dense model with the same total parameter count, but its overall scale still creates a significant infrastructure and latency burden.

Main strengths and limitations

Strengths

  • Open-weight control: teams can download, inspect, fine-tune, and self-host the checkpoint under its modified MIT license.
  • Large context: the verified 128K-token context window is useful for long documents, codebases, and multi-document workflows.
  • Technical focus: the model is positioned for coding, knowledge, reasoning, research, and custom language-system development.
  • Flexible deployment: official materials reference several established inference and serving stacks.
  • Foundation-model role: developers can adapt the model instead of accepting the behavior and restrictions of a finished assistant.

Limitations

  • Very large hardware footprint: the approximately 1.03 TB repository and trillion-parameter architecture are impractical for many local environments.
  • Not turnkey chat: the Base checkpoint may require instruction tuning, careful prompting, safety controls, and validation before end-user use.
  • Limited verified platform features: native tool use, JSON mode, structured output, caching, batch processing, and a hosted API are not confirmed for this exact model.
  • Text only: no native image, audio, or video input or output is documented for Kimi K2 Base.
  • Speed and cost trade-offs: deployment is more demanding than typical smaller open-weight models, and the supplied editorial speed assessment is low.
  • No verified output maximum: although the context window is documented, a separate maximum output-token limit is not.

When to choose Kimi K2 Base

Choose Kimi K2 Base when the central requirement is a large, adaptable language foundation rather than an immediately usable chat product. It is a reasonable candidate for an organization that has multi-GPU infrastructure, wants to fine-tune a model for a specialized domain, needs self-hosting, or is conducting research on large mixture-of-experts systems.

It is also suitable when a 128K-token context window is valuable and the team can accept the engineering work required to turn a pretrained checkpoint into a dependable application. Coding research, internal language systems, long-context experiments, and custom evaluation pipelines are stronger fits than casual personal use.

Another option is more appropriate when the priority is low latency, low infrastructure cost, native multimodal input, managed API access, or reliable instruction-following without substantial post-training. Within the Kimi K2 family, Kimi K2 Instruct is the more relevant comparison for ready-to-use chat and agentic behavior. Smaller open-weight models may be preferable when hardware availability and response speed matter more than maximum model scale.

Bottom line

Kimi K2 Base is best understood as a large open-weight building block. Its 1 trillion total parameters, approximately 32 billion activated parameters per token, 128K context window, and modified MIT license make it interesting for serious customization and self-hosted research. Those same characteristics make it a demanding choice: the checkpoint is enormous, inference is not low-cost or low-latency, and the supplied research does not verify a finished hosted API or a complete set of assistant features.

For teams prepared to operate large-model infrastructure and build the missing application layers, Kimi K2 Base offers substantial flexibility. For users who simply need a fast conversational assistant or a managed multimodal service, a post-trained or smaller hosted model is likely to be a better fit.


Answers to Frequently Asked Questions

Who should choose Kimi K2 Base instead of Kimi K2 Instruct?
Kimi K2 Base is better suited to organizations with substantial GPU infrastructure that need maximum control, fine-tuning, self-hosting, or long-context research. Kimi K2 Instruct is generally more appropriate for ready-to-use chat and agentic interactions, while smaller models may be preferable when low cost and low latency are priorities.
Is Kimi K2 Base suitable for chat, tool use, or multimodal applications?
Kimi K2 Base is not presented as a turnkey chat assistant. Native tool use, function calling, JSON mode, structured outputs, and multimodal image, audio, or video support are not verified for this exact checkpoint. Developers would need to add instruction tuning, orchestration, validation, and safety controls.
Can Kimi K2 Base be fine-tuned and self-hosted?
Yes. Kimi K2 Base is distributed as downloadable open weights under a modified MIT license and can be adapted for specialized domains, coding systems, research workflows, and self-hosted applications. Deployment may require infrastructure such as Transformers, vLLM, SGLang, KTransformers, or TensorRT-LLM-compatible systems.
What is Kimi K2 Base?
Kimi K2 Base is Moonshot AI’s open-weight pretrained foundation model in the Kimi K2 family. It is designed for self-hosting, fine-tuning, research, and custom language applications rather than immediate use as a finished chat assistant.
How large is Kimi K2 Base and what context window does it support?
Kimi K2 Base has 1 trillion total parameters, approximately 32 billion activated parameters per token, and a 131,072-token context window commonly described as 128K. Its repository is approximately 1.03 TB, so it requires substantial storage and multi-GPU infrastructure.


Sources 5
Provider

About Moonshot AI