Kimi K2

Kimi K2 Thinking

by Moonshot AI · Retired from Moonshot's direct API; open-weight checkpoint remains available for self-hosted or third-party deployment

An overview of Kimi K2 Thinking, Moonshot AI’s open-weight reasoning model for long-horizon research, coding, and tool-using agents. It covers the 1T-parameter mixture-of-experts architecture, 256K context window, text-only modality, deployment options, limitations, unavailable direct API pricing, and when another model may be a better fit.

Text Reasoning Coding
Kimi K2 Thinking is an open-weight language model released by Moonshot AI on November 6, 2025. It was designed for tasks that require more than a single response: planning, inspecting tool results, revising a plan, writing code, researching information, and continuing through many sequential actions. The model combines extended reasoning with native function calling and supports a 256K-token context window. Its main practical limitation is availability: Moonshot’s current catalog marks both kimi-k2-thinking and kimi-k2-thinking-turbo as offline, with the Kimi K2 series retired from the direct API on May 25, 2026.
Outputs

What Kimi K2 Thinking can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Kimi K2
Model type Reasoning
Context window 262K tokens
Release date 2025-11-06
Status Retired from Moonshot's direct API; open-weight checkpoint remains available for self-hosted or third-party deployment
Deprecation date 2026-05-25
Shutdown date 2026-05-25
Knowledge cutoff notes

Moonshot's official release material and model card do not provide a verifiable knowledge-cutoff date for Kimi K2 Thinking.

Model notes

Moonshot announced Kimi K2 Thinking on November 6, 2025. The current Kimi API model catalog lists kimi-k2-thinking and kimi-k2-thinking-turbo as offline and states that the Kimi K2 series was taken offline on May 25, 2026. The downloadable checkpoint is moonshotai/Kimi-K2-Thinking. The model card describes a 1T total-parameter MoE architecture with 32B activated parameters, 384 experts, eight selected experts per token, and a 256K context window. It supports text generation, reasoning output, and tool orchestration. Direct Moonshot pricing is not current because the API identity is retired; third-party hosted prices should not be treated as Moonshot list pricing.

Model guide

Kimi K2 Thinking: Open-Weight Reasoning for Long-Horizon Agents

Kimi K2 Thinking is Moonshot AI’s open-weight reasoning model for extended analysis, coding, research, and multi-step tool use. Its 1-trillion-parameter mixture-of-experts design, 256K-token context window, and ability to interleave reasoning with tool calls make it suited to long agentic workflows. However, the direct Moonshot API model was taken offline on May 25, 2026, so practical use now generally means self-hosting or using an independent hosted deployment.

What is Kimi K2 Thinking?

Kimi K2 Thinking is a reasoning-focused large language model from Moonshot AI, the provider behind the Kimi family. It is distributed as downloadable model weights rather than only as a hosted chat feature. The official checkpoint is named moonshotai/Kimi-K2-Thinking, and the released model and associated code use a modified MIT license.

The model is intended for long-running tasks in which the system must reason, use tools, examine their results, and decide what to do next. Typical examples include autonomous research, software development, complex information gathering, and multi-step writing. Unlike a simple text-generation model that produces one answer from one prompt, Kimi K2 Thinking was designed to alternate between internal reasoning and external actions during a single workflow.

Moonshot announced the model on November 6, 2025. Its direct API identity is no longer active: the current Kimi API catalog lists kimi-k2-thinking and kimi-k2-thinking-turbo as offline and records May 25, 2026 as the date the Kimi K2 series was taken offline. The weights remain available for self-hosted or third-party deployment, but those options should not be confused with a currently supported first-party Moonshot endpoint.

Architecture and 256K context window

Kimi K2 Thinking uses a mixture-of-experts, or MoE, architecture. It has approximately 1 trillion total parameters, but about 32 billion parameters are activated for each token. In practical terms, the model has a very large overall capacity without using every parameter for every piece of text. This can support a large model design while making inference different from a conventional dense model of the same total size.

The model card specifies 61 layers, 384 experts, eight selected experts per token, a 160K vocabulary, and MLA attention. Its context length is 256K tokens, or 262,144 tokens in the supplied model metadata. A context window is the amount of text and other conversation state the model can consider in one request. The large window is particularly relevant to source-heavy research, large code repositories, long documents, and agent sessions that accumulate tool results over time.

The supplied research does not specify a maximum output-token limit. It is therefore safer to treat the 256K figure as the context limit, not as a guarantee that a single response can contain 256K newly generated tokens.

Reasoning and tool use

The defining capability of Kimi K2 Thinking is its combination of extended reasoning and tool orchestration. Moonshot describes the model as able to interleave reasoning with tool calls, allowing an application to give it functions such as search, code execution, file operations, or other external actions. The model can then plan an action, inspect the returned result, and continue reasoning from that result.

Moonshot reported that the model could maintain workflows involving approximately 200 to 300 sequential tool invocations. This is a provider claim rather than an independently verified guarantee, and real performance will depend on the application’s tool definitions, error handling, context management, and inference setup. Nevertheless, the reported range illustrates the intended use case: Kimi K2 Thinking is aimed at extended agent loops rather than only short question-and-answer exchanges.

Compatible Moonshot interfaces can return reasoning in a separate reasoning_content field. Applications that use tool calling may need to preserve the complete assistant message, including the relevant reasoning and tool-call structure, when sending a follow-up request. The exact handling depends on the inference framework or hosted provider being used.

Coding, research, and writing performance

Kimi K2 Thinking is positioned for coding tasks that involve planning and iteration rather than simple code completion. An agent can inspect a project, call development tools, interpret test output, modify files, and continue through additional checks. This makes the model a plausible fit for repository analysis, debugging workflows, code generation, and autonomous programming agents.

For research, the long context and tool-oriented design can help an application collect multiple documents or results before producing a synthesis. The model can also be used for complex information gathering in which each action depends on what was found previously. It should not be assumed to have built-in web access, however. The supplied model metadata does not verify a native web-search capability for the model itself; web browsing or search must be provided by the surrounding application when supported.

The model also supports general text generation for analysis and writing. Its strength is most relevant when the assignment involves several stages, competing sources, structured investigation, or repeated revisions. For a short email, basic summarization, or a fast conversational response, a smaller or faster model may be more practical.

Supported modalities and output types

Kimi K2 Thinking is a text model. The supplied specifications identify text input and text output, with no verified image, audio, or video input. It does not provide native image, audio, or video generation. Its reasoning output is still text, even when it is exposed separately from the final answer.

Function calling and tool use extend what an application can accomplish, but they do not make the underlying model natively multimodal. An agent may process images or files if the surrounding system converts them into supported representations or supplies appropriate tools, but that should not be described as native image understanding without separate evidence.

Deployment and availability

The official weights can be deployed with inference engines such as vLLM, SGLang, and KTransformers. Because the model has approximately 1 trillion total parameters, self-hosting requires substantial accelerator memory and distributed-inference infrastructure. The model is therefore not a lightweight local download for an ordinary laptop or a small single-GPU setup.

Independent hosted services may make the checkpoint available, but their pricing, latency, quotas, privacy policies, and supported API formats are separate from Moonshot’s former API offering. Users should verify the provider’s exact checkpoint, quantization, context limit, tool-calling behavior, and data policy before relying on such a service.

There is no current first-party Moonshot price to report for Kimi K2 Thinking. The direct model identity is retired, and the supplied research does not provide a current self-hosting or third-party price. Any advertised hosted price should be treated as the price of that particular deployment, not as an official Moonshot list price for the retired API model.

Main strengths and limitations

Strengths

  • Long context: The 256K-token context window is suitable for large documents, codebases, and lengthy agent histories.
  • Extended reasoning: The model is designed for deliberate multi-step analysis rather than only immediate responses.
  • Native tool orchestration: Function calling and interleaved reasoning support agents that repeatedly act and inspect results.
  • Open-weight access: Developers can download the checkpoint and choose supported self-hosted or independent hosting arrangements.
  • Broad text use: The model covers research, coding, analysis, and writing within one text-generation system.

Limitations

  • First-party API retirement: Moonshot’s direct Kimi K2 Thinking API endpoint is offline as of May 25, 2026.
  • Heavy deployment requirements: The model’s scale makes self-hosting technically and financially demanding.
  • No native non-text generation: It does not generate images, audio, or video and has no verified native image, audio, or video input.
  • Tool infrastructure is external: Search, browsing, code execution, and file actions require an application or hosting layer that supplies those tools.
  • Unspecified output ceiling: The supplied documentation confirms the context window but does not provide a verified maximum output-token limit.
  • Deployment variability: Third-party services may differ in quantization, latency, limits, model version, and tool support.

When to choose Kimi K2 Thinking

Choose Kimi K2 Thinking when you need an open-weight reasoning model for a long-running agent and are prepared to manage substantial infrastructure or evaluate an independent hosted deployment. It is especially appropriate for autonomous research, complex coding agents, document-heavy analysis, and workflows that may require dozens or hundreds of tool interactions.

The model is also a reasonable choice when control over deployment matters more than the convenience of a current managed API. Its downloadable weights allow teams to investigate self-hosting, customize serving infrastructure, and avoid depending exclusively on the retired Moonshot endpoint.

Another option may be more appropriate when fast responses, low operating cost, or simple integration are priorities. A smaller dense model can be easier to run and may be preferable for routine classification, short answers, and basic coding assistance. A currently supported hosted model is a better fit for teams that do not want to maintain distributed inference. A multimodal model should be selected for native image understanding or image, audio, or video generation, since Kimi K2 Thinking is text-only.

Bottom line

Kimi K2 Thinking is best understood as an open-weight, long-context reasoning engine for agent builders rather than as a currently available general-purpose Moonshot API model. Its combination of a 256K context window, extended reasoning, and repeated tool use is well matched to complex workflows that need planning and iteration. The trade-off is substantial deployment complexity, no native non-text modalities, unspecified maximum output length, and the retirement of its direct first-party API. For teams able to host or independently access the checkpoint, it remains a specialized option for long-horizon text agents; for users seeking a simple, current, low-cost API, another model type is likely more suitable.


Answers to Frequently Asked Questions

What are the main limitations of Kimi K2 Thinking?
The model requires substantial infrastructure because it has approximately 1 trillion total parameters, making self-hosting difficult for ordinary laptops or small single-GPU systems. It is text-only, has no verified native image, audio, or video capabilities, requires external tools for browsing and code execution, has an unspecified maximum output length, and no longer has an active first-party Moonshot API.
Can Kimi K2 Thinking use tools and perform autonomous tasks?
Yes. Kimi K2 Thinking is designed to interleave reasoning with tool calls, allowing applications to provide functions such as search, code execution, file operations, and other external actions. Moonshot reported workflows involving approximately 200 to 300 sequential tool invocations, although actual performance depends on the tools, infrastructure, context management, and error handling.
Is the Kimi K2 Thinking API still available from Moonshot AI?
No. Moonshot’s direct API identities for Kimi K2 Thinking are listed as offline, with the Kimi K2 series taken offline on May 25, 2026. The model weights remain available for self-hosting or deployment through independent providers, but those options are separate from a currently supported first-party Moonshot endpoint.
What is Kimi K2 Thinking?
Kimi K2 Thinking is an open-weight, reasoning-focused large language model from Moonshot AI. It is designed for long-running agent workflows that combine multi-step reasoning with tool use, including autonomous research, software development, information gathering, and complex writing.
What is the context window of Kimi K2 Thinking?
Kimi K2 Thinking has a 256K-token context window, specified as 262,144 tokens in its model metadata. This makes it suitable for large documents, code repositories, source-heavy research, and agent sessions with extensive tool results. The available documentation does not confirm a maximum output-token limit.


Sources 5
Provider

About Moonshot AI