Kimi K2

Kimi K2.7 Code

by Moonshot AI · Current and accessible through Kimi API; Kimi Code default service has moved to Kimi K2.8 Preview, while Kimi K2.7 Code remains available through API and Kimi K2.7 Code HighSpeed service.

Kimi K2.7 Code is Moonshot AI’s open-weight coding and agent model for repository-scale software engineering. It provides a 256K-token context window, mandatory thinking, tool calling, image input, experimental API video input, and hosted or local deployment options. Its main trade-off is the substantial cost and infrastructure required for long-horizon reasoning and local inference.

Text Reasoning Coding
Kimi K2.7 Code is a specialized coding model from Moonshot AI for long-horizon software engineering. Released on June 12, 2026, it is designed to maintain context across large repositories, tool calls, debugging sessions, and multi-step coding tasks. The model is available through the Kimi API, Kimi Code services, and an open-weight release, but its large hardware requirements make hosted access more practical for many users.
Outputs

What Kimi K2.7 Code can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Streaming Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Kimi K2
Model type Coding
Context window 262K tokens
Release date 2026-06-12
Status Current and accessible through Kimi API; Kimi Code default service has moved to Kimi K2.8 Preview, while Kimi K2.7 Code remains available through API and Kimi K2.7 Code HighSpeed service.
Knowledge cutoff notes

Moonshot AI's current model documentation and model card do not publish a directly verifiable knowledge-cutoff date for Kimi K2.7 Code.

Model notes

Kimi K2.7 Code is a thinking-only model and forces preserve_thinking across multi-turn interactions. It has approximately 1T total parameters and 32B activated parameters, with a 400M MoonViT vision encoder. Official API video input is experimental and is not equivalently supported by third-party vLLM or SGLang deployments. The open-weight checkpoint is available on Hugging Face under a modified MIT license and uses native INT4 quantization. Kimi Code uses kimi-for-coding as a service/model identity, but current Kimi Code documentation states that this identity now routes to Kimi K2.8 Preview; Kimi K2.7 Code remains a distinct model available through the Kimi API and the Kimi K2.7 Code HighSpeed service. Editorial scores are comparative estimates rather than provider-published ratings.

Cost

Model pricing

Input ¥6.50 per 1M uncached input tokens; ¥1.30 per 1M cache-hit input tokens
Output ¥27.00 per 1M output tokens
Model guide

Kimi K2.7 Code: Moonshot AI’s Open-Weight Agent for Long-Horizon Software Engineering

Kimi K2.7 Code is Moonshot AI’s open-weight, coding-focused agentic model for complex software engineering tasks. It combines mandatory reasoning, tool calling, native image input, experimental video input, a 256K-token context window, and a 1-trillion-parameter mixture-of-experts architecture with 32 billion active parameters. It is best suited to repository-scale coding, debugging, refactoring, and multi-step development workflows rather than general conversation or lightweight local deployment.

What is Kimi K2.7 Code?

Kimi K2.7 Code is Moonshot AI’s coding-focused reasoning and agent model. Its purpose is not simply to autocomplete short code snippets. It is designed for longer software engineering workflows in which the model must inspect a repository, reason about dependencies, call tools, modify multiple files, run or interpret development tasks, and continue working across several steps.

The model is based on the Kimi K2.6 architecture and operates in thinking mode only. In practical terms, every request uses the model’s reasoning workflow; there is no supported instant or non-thinking mode. Moonshot AI also describes Kimi K2.7 Code as requiring the preservation of reasoning content across multi-turn interactions, which is intended to help it maintain continuity during coding-agent tasks.

Kimi K2.7 Code was released on June 12, 2026. It is available through the hosted Kimi API under the model identifier kimi-k2.7-code, through relevant Kimi Code services, and as an open-weight checkpoint on Hugging Face under a modified MIT license.

Where it fits in Moonshot AI’s current lineup

Kimi K2.7 Code occupies the specialized coding and agentic-development position in Moonshot AI’s model lineup. It should not be treated as a general-purpose replacement for the broader Kimi conversational models. Moonshot AI recommends Kimi K2.6 for writing, analysis, and general dialogue, while Kimi K2.7 Code is aimed at software engineering and tool-driven development.

There is also an important distinction between the model and the current Kimi Code default service. As of September 25, 2026, Kimi Code documentation indicates that the default kimi-for-coding service has moved to Kimi K2.8 Preview. Kimi K2.7 Code nevertheless remains separately accessible through the Kimi API and the Kimi K2.7 Code HighSpeed service. This means users should check the selected model or service identity rather than assume that every Kimi Code request uses K2.7 Code.

Architecture and 256K context window

Kimi K2.7 Code uses a mixture-of-experts, or MoE, architecture. It has approximately 1 trillion total parameters, but about 32 billion parameters are activated for each token. This routing approach allows the model to have a very large overall network while using only a subset of experts for each part of a request.

The documented architecture has 61 layers, 384 experts, and eight experts selected per token. It also includes Multi-head Latent Attention and a 400-million-parameter MoonViT vision encoder for visual input.

The context window is 256K tokens, or 262,144 tokens. A context window is the amount of text, code, reasoning history, and tool information the model can consider in one interaction. This capacity is useful for large repositories, extensive logs, long debugging sessions, and agent workflows that need to preserve substantial tool history. It does not mean that every repository can be inserted without preparation: practical usage still depends on how the codebase is selected, retrieved, and organized.

The supplied documentation does not publish a maximum output-token limit for Kimi K2.7 Code. Users should therefore avoid assuming that the 256K context window represents the maximum length of a single generated answer; context capacity and output capacity are separate limits.

Coding, reasoning, and agent capabilities

The model’s central strength is its combination of coding ability and persistent reasoning. It is intended for tasks such as:

  • Repository-level feature development involving multiple files.
  • Large-scale refactoring where changes must remain consistent across modules.
  • Debugging that requires reading error output, tracing dependencies, and revising an implementation.
  • Tool-driven coding workflows and agent frameworks.
  • Long-running software engineering tasks that require several intermediate decisions.

Tool calling is supported, allowing an integrating application or coding environment to expose functions to the model. Those functions may let an agent inspect files, search a codebase, invoke development tools, or perform other controlled actions. The model itself does not make external actions automatically in every deployment; the surrounding application determines which tools are available and what permissions they have.

Moonshot AI reports improvements over Kimi K2.6 on coding and agent benchmarks including Kimi Code Bench v2, Program Bench, MLS Bench Lite, Kimi Claw 24/7 Bench, MCP Atlas, and MCP Mark Verified. These are provider-reported benchmark claims rather than independent evaluations. Moonshot AI also states that Kimi K2.7 Code uses approximately 30% fewer thinking tokens than Kimi K2.6, which may improve efficiency in comparable workloads, although actual latency and cost still depend on request length, caching, infrastructure, and service configuration.

Supported input and output modalities

Kimi K2.7 Code accepts text and image input. Image understanding can be useful when a coding task includes screenshots of an interface, diagrams, visual bugs, or other visual references. The official API also supports video input, but Moonshot AI describes that capability as experimental.

Video support is not equivalent across all deployment methods. The supplied model documentation notes that official API video input is not currently supported in the same way by third-party vLLM or SGLang deployments. Users planning to run the open-weight model locally should therefore distinguish between the model’s documented modality support and the features implemented by a particular inference engine.

The model produces text and reasoning output. It does not natively generate images, video, audio, music, or other media. Its multimodal capability is therefore primarily an input capability, not a creative-output capability.

Pricing and access options

Hosted Kimi API pricing is listed at ¥6.50 per million uncached input tokens, ¥1.30 per million cache-hit input tokens, and ¥27.00 per million output tokens. The API supports automatic context caching. Cache-hit pricing can reduce the cost of repeated context, such as a persistent repository instruction set or recurring tool history, when the request pattern qualifies for caching.

The input and output prices apply to different parts of a request. Sending a large codebase can create substantial input-token charges, while a long generated patch or reasoning-heavy response contributes to output charges. Thinking-only operation may also make the model less suitable for very small tasks where a quick response matters more than extended reasoning.

The open-weight checkpoint provides another access route, but local operation is demanding. The native INT4 checkpoint is approximately 595 GB on disk, and practical inference requires substantial multi-GPU infrastructure. Moonshot AI recommends inference engines including vLLM, SGLang, and KTransformers. Open weights therefore provide deployment flexibility without making the model lightweight or inexpensive to host.

Main strengths and limitations

Strengths:

  • A 256K-token context window suited to large codebases and long agent sessions.
  • Purpose-built support for repository-scale programming and multi-step software engineering.
  • Mandatory reasoning for tasks that benefit from deliberate planning and debugging.
  • Tool calling and multi-step workflow support.
  • Text and image input, plus experimental video input through the official API.
  • Hosted API access alongside an open-weight release.
  • Lower reported thinking-token usage than Kimi K2.6.

Limitations:

  • There is no supported instant or non-thinking mode.
  • The model is very large to deploy locally despite its 32-billion-parameter active routing size.
  • The approximately 595 GB INT4 checkpoint requires substantial multi-GPU infrastructure.
  • Video input is experimental and deployment-dependent.
  • It does not generate images, video, audio, or music.
  • The supplied documentation does not specify a maximum output-token limit.
  • It is specialized for coding and is not the recommended option for general writing, analysis, or casual conversation.
  • Current Kimi Code defaults may route to Kimi K2.8 Preview rather than Kimi K2.7 Code.

When to choose Kimi K2.7 Code

Choose Kimi K2.7 Code when the task involves substantial software context and benefits from deliberate, multi-step reasoning. It is a strong fit for repository analysis, complex refactoring, debugging across several files, codebase migration work, and agents that need to combine model reasoning with tools.

Its large context window is especially useful when repeatedly supplying architecture notes, source files, test results, and tool outputs would otherwise force the application to discard important history. The hosted API is generally the more practical choice for teams that want access without operating a very large multi-GPU deployment.

Another coding model may be more appropriate for short code completions, low-latency interactive editing, or cost-sensitive requests where mandatory reasoning adds unnecessary overhead. Kimi K2.6 is the better-positioned Moonshot AI option when the primary task is general writing, analysis, or dialogue. Users who need media generation should also choose a model or product designed to produce those output types, because Kimi K2.7 Code only returns text and reasoning content.

Overall assessment

Kimi K2.7 Code is best understood as a large, specialized coding-agent model rather than a general conversational assistant. Its defining combination is a 256K-token context window, thinking-only operation, tool support, multimodal input, and an open-weight architecture designed for long-horizon engineering tasks.

The trade-off is substantial resource use. Hosted access avoids the model’s extreme local hardware requirements, but token costs and reasoning latency still matter. For teams building serious coding agents or handling repository-scale work, those costs may be justified by the model’s context and workflow capabilities. For quick edits, lightweight deployment, or ordinary conversation, a smaller or more general model is likely to be a better fit.


Answers to Frequently Asked Questions

Is Kimi K2.7 Code the default model in Kimi Code?
Not necessarily. As of September 25, 2026, the default kimi-for-coding service uses Kimi K2.8 Preview. Kimi K2.7 Code remains available separately through the Kimi API and the Kimi K2.7 Code HighSpeed service, so users should verify the selected model or service.
Can Kimi K2.7 Code be run locally?
Yes, Kimi K2.7 Code is available as an open-weight checkpoint, but local deployment requires substantial infrastructure. The native INT4 checkpoint is approximately 595 GB and practical inference requires multi-GPU hardware. Moonshot AI recommends engines such as vLLM, SGLang, and KTransformers.
How much does the hosted Kimi K2.7 Code API cost?
The hosted API pricing is ¥6.50 per million uncached input tokens, ¥1.30 per million cache-hit input tokens, and ¥27.00 per million output tokens. Automatic context caching can reduce costs for repeated repository instructions or tool history when requests qualify.
What is the context window of Kimi K2.7 Code?
Kimi K2.7 Code has a 256K-token context window, equal to 262,144 tokens. This is useful for large repositories, lengthy debugging sessions, tool outputs, logs, and other workflows that require substantial context. The documentation does not specify a separate maximum output-token limit.
What is Kimi K2.7 Code designed for?
Kimi K2.7 Code is a coding-focused reasoning and agent model designed for long software engineering workflows. It can inspect repositories, reason about dependencies, call tools, modify multiple files, debug issues, and maintain context across multi-step tasks.


Sources 5
Provider

About Moonshot AI