What is Kimi K2.7 Code?
Kimi K2.7 Code is Moonshot AI’s coding-focused reasoning and agent model. Its purpose is not simply to autocomplete short code snippets. It is designed for longer software engineering workflows in which the model must inspect a repository, reason about dependencies, call tools, modify multiple files, run or interpret development tasks, and continue working across several steps.
The model is based on the Kimi K2.6 architecture and operates in thinking mode only. In practical terms, every request uses the model’s reasoning workflow; there is no supported instant or non-thinking mode. Moonshot AI also describes Kimi K2.7 Code as requiring the preservation of reasoning content across multi-turn interactions, which is intended to help it maintain continuity during coding-agent tasks.
Kimi K2.7 Code was released on June 12, 2026. It is available through the hosted Kimi API under the model identifier kimi-k2.7-code, through relevant Kimi Code services, and as an open-weight checkpoint on Hugging Face under a modified MIT license.
Where it fits in Moonshot AI’s current lineup
Kimi K2.7 Code occupies the specialized coding and agentic-development position in Moonshot AI’s model lineup. It should not be treated as a general-purpose replacement for the broader Kimi conversational models. Moonshot AI recommends Kimi K2.6 for writing, analysis, and general dialogue, while Kimi K2.7 Code is aimed at software engineering and tool-driven development.
There is also an important distinction between the model and the current Kimi Code default service. As of September 25, 2026, Kimi Code documentation indicates that the default kimi-for-coding service has moved to Kimi K2.8 Preview. Kimi K2.7 Code nevertheless remains separately accessible through the Kimi API and the Kimi K2.7 Code HighSpeed service. This means users should check the selected model or service identity rather than assume that every Kimi Code request uses K2.7 Code.
Architecture and 256K context window
Kimi K2.7 Code uses a mixture-of-experts, or MoE, architecture. It has approximately 1 trillion total parameters, but about 32 billion parameters are activated for each token. This routing approach allows the model to have a very large overall network while using only a subset of experts for each part of a request.
The documented architecture has 61 layers, 384 experts, and eight experts selected per token. It also includes Multi-head Latent Attention and a 400-million-parameter MoonViT vision encoder for visual input.
The context window is 256K tokens, or 262,144 tokens. A context window is the amount of text, code, reasoning history, and tool information the model can consider in one interaction. This capacity is useful for large repositories, extensive logs, long debugging sessions, and agent workflows that need to preserve substantial tool history. It does not mean that every repository can be inserted without preparation: practical usage still depends on how the codebase is selected, retrieved, and organized.
The supplied documentation does not publish a maximum output-token limit for Kimi K2.7 Code. Users should therefore avoid assuming that the 256K context window represents the maximum length of a single generated answer; context capacity and output capacity are separate limits.
Coding, reasoning, and agent capabilities
The model’s central strength is its combination of coding ability and persistent reasoning. It is intended for tasks such as:
- Repository-level feature development involving multiple files.
- Large-scale refactoring where changes must remain consistent across modules.
- Debugging that requires reading error output, tracing dependencies, and revising an implementation.
- Tool-driven coding workflows and agent frameworks.
- Long-running software engineering tasks that require several intermediate decisions.
Tool calling is supported, allowing an integrating application or coding environment to expose functions to the model. Those functions may let an agent inspect files, search a codebase, invoke development tools, or perform other controlled actions. The model itself does not make external actions automatically in every deployment; the surrounding application determines which tools are available and what permissions they have.
Moonshot AI reports improvements over Kimi K2.6 on coding and agent benchmarks including Kimi Code Bench v2, Program Bench, MLS Bench Lite, Kimi Claw 24/7 Bench, MCP Atlas, and MCP Mark Verified. These are provider-reported benchmark claims rather than independent evaluations. Moonshot AI also states that Kimi K2.7 Code uses approximately 30% fewer thinking tokens than Kimi K2.6, which may improve efficiency in comparable workloads, although actual latency and cost still depend on request length, caching, infrastructure, and service configuration.
Supported input and output modalities
Kimi K2.7 Code accepts text and image input. Image understanding can be useful when a coding task includes screenshots of an interface, diagrams, visual bugs, or other visual references. The official API also supports video input, but Moonshot AI describes that capability as experimental.
Video support is not equivalent across all deployment methods. The supplied model documentation notes that official API video input is not currently supported in the same way by third-party vLLM or SGLang deployments. Users planning to run the open-weight model locally should therefore distinguish between the model’s documented modality support and the features implemented by a particular inference engine.
The model produces text and reasoning output. It does not natively generate images, video, audio, music, or other media. Its multimodal capability is therefore primarily an input capability, not a creative-output capability.
Pricing and access options
Hosted Kimi API pricing is listed at ¥6.50 per million uncached input tokens, ¥1.30 per million cache-hit input tokens, and ¥27.00 per million output tokens. The API supports automatic context caching. Cache-hit pricing can reduce the cost of repeated context, such as a persistent repository instruction set or recurring tool history, when the request pattern qualifies for caching.
The input and output prices apply to different parts of a request. Sending a large codebase can create substantial input-token charges, while a long generated patch or reasoning-heavy response contributes to output charges. Thinking-only operation may also make the model less suitable for very small tasks where a quick response matters more than extended reasoning.
The open-weight checkpoint provides another access route, but local operation is demanding. The native INT4 checkpoint is approximately 595 GB on disk, and practical inference requires substantial multi-GPU infrastructure. Moonshot AI recommends inference engines including vLLM, SGLang, and KTransformers. Open weights therefore provide deployment flexibility without making the model lightweight or inexpensive to host.
Main strengths and limitations
Strengths:
- A 256K-token context window suited to large codebases and long agent sessions.
- Purpose-built support for repository-scale programming and multi-step software engineering.
- Mandatory reasoning for tasks that benefit from deliberate planning and debugging.
- Tool calling and multi-step workflow support.
- Text and image input, plus experimental video input through the official API.
- Hosted API access alongside an open-weight release.
- Lower reported thinking-token usage than Kimi K2.6.
Limitations:
- There is no supported instant or non-thinking mode.
- The model is very large to deploy locally despite its 32-billion-parameter active routing size.
- The approximately 595 GB INT4 checkpoint requires substantial multi-GPU infrastructure.
- Video input is experimental and deployment-dependent.
- It does not generate images, video, audio, or music.
- The supplied documentation does not specify a maximum output-token limit.
- It is specialized for coding and is not the recommended option for general writing, analysis, or casual conversation.
- Current Kimi Code defaults may route to Kimi K2.8 Preview rather than Kimi K2.7 Code.
When to choose Kimi K2.7 Code
Choose Kimi K2.7 Code when the task involves substantial software context and benefits from deliberate, multi-step reasoning. It is a strong fit for repository analysis, complex refactoring, debugging across several files, codebase migration work, and agents that need to combine model reasoning with tools.
Its large context window is especially useful when repeatedly supplying architecture notes, source files, test results, and tool outputs would otherwise force the application to discard important history. The hosted API is generally the more practical choice for teams that want access without operating a very large multi-GPU deployment.
Another coding model may be more appropriate for short code completions, low-latency interactive editing, or cost-sensitive requests where mandatory reasoning adds unnecessary overhead. Kimi K2.6 is the better-positioned Moonshot AI option when the primary task is general writing, analysis, or dialogue. Users who need media generation should also choose a model or product designed to produce those output types, because Kimi K2.7 Code only returns text and reasoning content.
Overall assessment
Kimi K2.7 Code is best understood as a large, specialized coding-agent model rather than a general conversational assistant. Its defining combination is a 256K-token context window, thinking-only operation, tool support, multimodal input, and an open-weight architecture designed for long-horizon engineering tasks.
The trade-off is substantial resource use. Hosted access avoids the model’s extreme local hardware requirements, but token costs and reasoning latency still matter. For teams building serious coding agents or handling repository-scale work, those costs may be justified by the model’s context and workflow capabilities. For quick edits, lightweight deployment, or ordinary conversation, a smaller or more general model is likely to be a better fit.

