Kimi-Dev

Kimi-Dev-72B

by Moonshot AI · Available open-weight model

Moonshot AI's Kimi-Dev-72B is an approximately 73-billion-parameter open-weight coding model trained for repository-level issue resolution, code repair, and test writing. It uses a documented 131,072-token serving length, is distributed under the MIT license, and supports deployment through Transformers, vLLM, and SGLang. Its main trade-off is high infrastructure cost: hosted pricing and a separate maximum output limit are not documented.

Text Reasoning Coding
Kimi-Dev-72B is an open-weight coding model from Moonshot AI built for practical software engineering. Instead of focusing mainly on short code completion, it is designed to inspect repositories, identify the files that need changes, edit code, and write tests for reported issues. The model is available through Hugging Face and can be deployed with Transformers, vLLM, or SGLang on suitable infrastructure.
Outputs

What Kimi-Dev-72B can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

7/10 Reasoning
9/10 Coding
4/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Kimi-Dev
Model type Coding
Context window 131K tokens
Release date June 2025
Status Available open-weight model
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was identified in the official model card or Moonshot AI repository.

Model notes

Kimi-Dev-72B is distributed as an open-weight BF16 checkpoint under the MIT license. Moonshot AI reports 60.4% on SWE-bench Verified and describes reinforcement-learning training in Docker environments where rewards depend on complete test-suite success. The model is based on Qwen2.5-72B according to the Hugging Face model tree. Official materials document a 131,072-token serving length, but do not specify a separate maximum output-token limit, knowledge cutoff, hosted API price, structured-output guarantee, prompt-caching feature, batch API, or provider-managed web-search tool. Tool execution depends on the deployment or agent framework.

Model guide

Kimi-Dev-72B: Open-Weight Coding Model for Repository-Level Software Repair

Kimi-Dev-72B is Moonshot AI's open-weight coding model for repository-level issue resolution, code repair, file localization, and test generation. Released in June 2025 under the MIT license, it is designed for software-engineering agents and self-hosted deployment rather than ordinary chat. The approximately 73-billion-parameter BF16 model has a documented 131,072-token serving length and a provider-reported 60.4% score on SWE-bench Verified.

What is Kimi-Dev-72B?

Kimi-Dev-72B is Moonshot AI's coding-focused large language model for software engineering tasks. It was released in June 2025 as an open-weight checkpoint and is distributed under the MIT license. The model is listed as approximately 73 billion parameters and is based on Qwen2.5-72B according to its Hugging Face model tree.

The important distinction is its target workflow. Kimi-Dev-72B is not primarily positioned as a general-purpose conversational assistant or a hosted coding chatbot. It is intended to work on repository-level tasks such as resolving reported issues, locating relevant files, modifying complete files, fixing bugs, and generating tests. That makes it most relevant to developers building coding agents or teams that want to run a coding model on their own infrastructure.

Moonshot AI reports a 60.4% result on SWE-bench Verified, a benchmark based on real software-engineering issues. This is a provider-reported benchmark claim, not a guarantee that the model will solve every repository task successfully. Results in practice depend on the repository, test quality, prompting, available tools, inference configuration, and the surrounding agent workflow.

How the model is trained for software engineering

Kimi-Dev-72B was optimized with large-scale reinforcement learning. In simple terms, the training process gives the model feedback based on whether its proposed repository changes satisfy the task's tests, rather than rewarding only locally plausible code.

Moonshot AI describes a Docker-based training and evaluation environment built around real software repositories. Docker provides a controlled environment in which generated changes can be applied and tested. The reported reward is tied to complete test-suite success, which encourages the model to account for interactions between files instead of treating a coding problem as an isolated completion.

The official workflow separates two useful stages of repository repair. First, the system identifies the files most likely to require changes. It then edits the relevant files and evaluates the result. Moonshot AI also provides rollout scripts for bug-fixing and test-writing tasks. This design is particularly suitable for an external agent that can inspect a repository, run commands, apply patches, and return test results to the model.

Capabilities, context, and supported modalities

Kimi-Dev-72B is a text-in, text-out coding model. The supplied model materials do not document native image, audio, or video input or output. It can generate code, explanations, patches, and tests as text, but it does not itself produce images, audio, or video.

SpecificationVerified information
ProviderMoonshot AI
ReleaseJune 2025
Model sizeApproximately 73 billion parameters
CheckpointBF16 open-weight model
LicenseMIT
Documented serving length131,072 tokens
Input and outputText input and text output
Hosted model priceNot documented in the supplied sources
Maximum output tokensNot separately specified

The 131,072-token figure is the documented sequence or serving length recommended in the deployment materials. The supplied sources do not define a separate maximum-output-token limit, so it should not be treated as a guaranteed generation length for every deployment. The effective limit can also depend on the serving framework, memory configuration, prompt size, and generation settings.

The long context is useful for supplying repository instructions, issue descriptions, source files, test files, and tool results in one workflow. However, a large context window does not remove the need for good file selection. Feeding an entire large repository into every request can increase memory use, latency, and the chance that important details are overlooked.

Coding and reasoning performance

Kimi-Dev-72B's main strength is multi-step coding work. A repository-level issue often requires more than producing a syntactically valid function: the system may need to understand the issue, trace behavior across files, identify the correct implementation, make a compatible change, and add or update tests. The model's training and rollout design are aimed at this sequence of actions.

Its reasoning is therefore best understood as software-engineering reasoning rather than a separately documented reasoning mode. The supplied materials do not identify a distinct reasoning budget, selectable reasoning level, or hidden chain-of-thought feature. The model can produce analysis and code through ordinary text generation, while the surrounding agent can provide the execution loop that makes the workflow useful.

For coding, the model is better matched to repository repair, issue resolution, test generation, and coding-agent research than to lightweight autocomplete. Moonshot AI's SWE-bench Verified result supports its positioning for issue-resolution tasks, but benchmark performance should be interpreted alongside the substantial hardware and serving requirements of a 72-billion-parameter model.

Deployment and infrastructure

Kimi-Dev-72B can be downloaded from Hugging Face and deployed with Transformers, vLLM, or SGLang. Moonshot AI's example uses an OpenAI-compatible local server and recommends a 131,072-token sequence length for vLLM serving. An OpenAI-compatible endpoint can make the model easier to connect to existing applications, but it does not turn the checkpoint into a Moonshot-hosted API product.

The BF16 checkpoint is large. In practice, deployment generally requires substantial GPU memory, multiple GPUs, quantization, or hosted inference infrastructure. The model is consequently a poor fit for an ordinary consumer laptop unless a suitable quantized deployment is available and its performance is acceptable. The supplied sources do not specify a single minimum GPU configuration, so hardware requirements should be estimated from the chosen framework, precision, context length, batch size, and concurrency rather than from parameter count alone.

Self-hosting offers control over the runtime, repository data, and integration with local tools. It also transfers responsibility for hardware, scaling, monitoring, security, model updates, and inference cost to the operator. No official hosted pricing for Kimi-Dev-72B is documented in the supplied research, so there is no verified per-token price to compare with commercial coding APIs.

Tools and agent workflows

Kimi-Dev-72B is suitable for use inside a coding agent, but tool execution is not an intrinsic output modality of the model. The supplied materials do not document a provider-managed web-search tool, native function-calling guarantee, or separate structured-output guarantee for this checkpoint.

An application can still place the model in a tool loop. For example, an external controller might ask it to inspect a directory, read selected files, propose a patch, run a test command, and revise the patch after seeing the test output. The controller, not the model alone, is responsible for executing shell commands, enforcing permissions, validating patches, and deciding which files or test results should be returned to the next prompt.

This distinction matters for production use. A text model can suggest a command or patch without having permission to run it. Any deployment that allows repository modification should use sandboxing, restricted credentials, approval steps, and validation against the project's tests.

Main strengths and limitations

Strengths

  • Repository-level focus: The model is designed around issue resolution and multi-file software changes rather than only isolated code completion.
  • Practical training objective: Docker-based reinforcement learning and test-based rewards target whether a complete software task works.
  • Open-weight access: The MIT license and downloadable checkpoint support private deployment, experimentation, and integration into custom coding-agent systems.
  • Large documented context: The 131,072-token serving length can accommodate substantial issue, code, and test context when infrastructure allows it.
  • Deployment flexibility: Transformers, vLLM, and SGLang provide multiple routes for local or hosted inference.

Limitations

  • High infrastructure cost: A roughly 73-billion-parameter BF16 checkpoint is demanding compared with small coding models and ordinary local software tools.
  • No verified hosted price: The supplied sources do not document a Moonshot-hosted API plan or model-specific input and output pricing.
  • Text-only model: Native image, audio, and video capabilities are not documented.
  • Undocumented advanced interfaces: There is no supplied verification of structured output, function calling, streaming, caching, batch API support, or fine-tuning for this model.
  • Tooling is external: Repository access, command execution, patch application, and test running require an agent framework or deployment layer.
  • No separate output limit is specified: The 131,072-token serving length should not be read as a guaranteed maximum generated response.

When to choose Kimi-Dev-72B

Choose Kimi-Dev-72B when the main requirement is an open-weight coding model for serious repository work and you can provide the infrastructure needed to serve it. It is a reasonable candidate for:

  • Self-hosted coding agents that must work with private repositories.
  • Automated investigation of software issues and multi-file bug fixes.
  • Test generation and repair workflows driven by an external execution loop.
  • Research into reinforcement learning for software-engineering agents.
  • Teams that prefer an MIT-licensed checkpoint over a closed hosted coding service.

A smaller coding model may be more appropriate when response speed, low hardware cost, or laptop deployment matters more than repository-level capability. A hosted coding API may be preferable when the team does not want to operate GPUs, serving infrastructure, security controls, and scaling. A multimodal model is a better choice when the task requires interpreting screenshots, diagrams, audio, or video, because those capabilities are not documented for Kimi-Dev-72B.

Within its intended category, Kimi-Dev-72B trades operational simplicity and speed for an open, high-capacity model aimed at difficult software tasks. The best choice depends on whether the value of private deployment and repository-level performance outweighs the cost and complexity of running a large checkpoint.

Bottom line

Kimi-Dev-72B is a specialized open-weight coding model, not a general consumer assistant or a ready-made Moonshot API plan. Its value comes from combining a large model, long documented serving length, repository-oriented training, and a permissive MIT license. The provider-reported SWE-bench Verified result makes it a notable candidate for software-engineering agents, while the lack of documented hosted pricing and the heavy deployment requirements make it most practical for teams with dedicated inference infrastructure or a suitable hosted deployment partner.


Answers to Frequently Asked Questions

Does Kimi-Dev-72B support tools, images, or function calling natively?
Kimi-Dev-72B is documented as a text-in, text-out model, with no verified native image, audio, or video support. Repository access, command execution, patch application, and testing must be provided by an external agent or deployment layer, and the supplied materials do not verify native function calling or structured-output guarantees.
How can Kimi-Dev-72B be deployed?
The model can be downloaded from Hugging Face and deployed with Transformers, vLLM, or SGLang. It can also be exposed through an OpenAI-compatible local server, although it is not a Moonshot-hosted API product.
What are the hardware requirements for running Kimi-Dev-72B?
Kimi-Dev-72B is a large BF16 model with approximately 73 billion parameters, so deployment generally requires substantial GPU memory, multiple GPUs, quantization, or hosted inference infrastructure. Exact requirements depend on the serving framework, context length, batch size, precision, and concurrency.
What is Kimi-Dev-72B designed for?
Kimi-Dev-72B is an open-weight coding model from Moonshot AI designed for repository-level software engineering tasks, including issue resolution, multi-file bug fixes, code changes, and test generation.
What license does Kimi-Dev-72B use?
Kimi-Dev-72B is distributed under the MIT license, which supports private deployment, experimentation, and integration into custom coding-agent systems.


Sources 4
Provider

About Moonshot AI