Kimi K3

Kimi K3

by Moonshot AI · Current and available; open-weight model

Kimi K3 is Moonshot AI’s open-weight multimodal model for long-context reasoning, software engineering, technical research, and agentic workflows. It accepts text, images, and video, supports up to one million tokens of context, returns text, offers tool calling and structured responses, and is priced through the international API at $3.00 per million cache-miss input tokens, $0.30 per million cache-hit input tokens, and $15.00 per million output tokens.

Text Reasoning Coding
Kimi K3 is Moonshot AI’s largest current model, released on July 16, 2026, for demanding coding, research, and agentic tasks. It accepts text, images, and video, processes up to one million tokens of context, returns text, and is available through Kimi products, the Kimi API, Kimi Code, and downloadable open weights.
Outputs

What Kimi K3 can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Fine-tuning Structured output Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
5/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Kimi K3
Model type Multimodal
Context window 1M tokens
Maximum output 131K tokens
Release date 2026-07-16
Status Current and available; open-weight model
Knowledge cutoff notes

Moonshot AI's official Kimi K3 model page, launch material, technical report, and API materials reviewed for this record do not state a calendar knowledge-cutoff date. Web search and retrieval features do not change the underlying cutoff.

Model notes

Kimi K3 is a 2.8-trillion-parameter sparse Mixture-of-Experts model with 104 billion activated parameters according to the technical report, using Kimi Delta Attention, Attention Residuals, Stable LatentMoE, and MoonViT-V2. The model accepts text, image, and video inputs and returns text. Moonshot AI provides the model through Kimi.com, Kimi Work, Kimi Code, and the Kimi API, and released downloadable weights and a technical report on July 27, 2026 under the Kimi K3 License. The official international API launch page lists $0.30 per MTok for cache-hit input, $3.00 per MTok for cache-miss input, and $15.00 per MTok for output. A separate Chinese platform lists regional CNY prices. The 131,072-token output limit is reported in current Kimi model catalogs and ecosystem documentation, while the context window is officially documented as up to one million tokens. No authoritative model-specific knowledge-cutoff date was found. Reasoning and coding scores are editorial comparative estimates rather than provider-published ratings.

Cost

Model pricing

Input $3.00 per MTok cache-miss input; $0.30 per MTok cache-hit input
Output $15.00 per MTok
Model guide

Kimi K3: Moonshot AI’s Open-Weight Model for Long-Context Agentic Coding

Kimi K3 is Moonshot AI’s flagship open-weight multimodal model for long-context reasoning, software engineering, technical research, and agentic workflows. It combines a sparse 2.8-trillion-parameter Mixture-of-Experts architecture with native vision, video understanding, a context window of up to one million tokens, tool use, structured outputs, and a maximum reported output of 131,072 tokens.

What is Kimi K3?

Kimi K3 is an open-weight multimodal model from Moonshot AI. It is designed for tasks that require extended reasoning over large amounts of information rather than for simple, low-latency chat. Its intended uses include software engineering, long-document analysis, technical research, tool-assisted work, and autonomous or semi-autonomous agents.

The model was released on July 16, 2026, and sits at the high-capability end of Moonshot AI’s current Kimi lineup. It is available through Kimi.com and related Kimi products, Kimi Work, Kimi Code, the Kimi API, and downloadable model weights. The availability of a downloadable version makes Kimi K3 relevant both to hosted users and to organizations evaluating open-weight deployment, although its size creates substantial infrastructure requirements.

Moonshot AI describes Kimi K3 as a model for long-horizon work: tasks in which the system must maintain context, reason through multiple steps, call tools, write and test code, and revise its approach. These are provider positioning claims, while the architecture and interface specifications below come from the supplied model documentation and catalog information.

Architecture and one-million-token context

Kimi K3 has 2.8 trillion total parameters and uses a sparse Mixture-of-Experts, or MoE, architecture. In an MoE model, only a subset of the available expert networks is activated for each token instead of running every parameter on every step. The supplied technical information reports 896 routed experts, with 16 activated per token, and approximately 104 billion activated parameters.

This design gives Kimi K3 a very large total capacity without requiring every expert to run for every token. It does not make the model easy to host on ordinary hardware, however. The total parameter count and recommended accelerator configurations make self-hosting considerably more demanding than deploying a smaller open-weight model.

The documented context window is up to 1,000,000 tokens. Context is the working information available to the model during a request, including instructions, conversation history, source documents, code, and multimodal content. A one-million-token limit can be useful for reviewing large codebases, lengthy technical archives, multiple documents, or extended agent sessions without repeatedly discarding earlier material.

A large context window is not a guarantee that every detail will receive equal attention, and practical performance will depend on the task, prompt structure, serving configuration, and available memory. It also does not provide a verified knowledge-cutoff date. Moonshot AI’s available Kimi K3 documentation does not publish a calendar cutoff for the model.

Inputs, outputs, and supported modalities

Kimi K3 accepts text, images, and video as inputs and produces text as its model output. Its native vision capability allows it to interpret visual material alongside written instructions. Practical examples include analyzing screenshots during software debugging, extracting information from visual documents, inspecting diagrams, and reviewing video content with a textual question.

The model does not natively generate images, audio, or video. This distinction matters because Kimi’s wider product ecosystem supports creative functions through supported products or plugins, but those product-level capabilities should not be attributed to Kimi K3 itself. Kimi K3 is a text-output reasoning model with multimodal understanding.

The reported maximum output is 131,072 tokens. That is a substantial allowance for long code revisions, detailed technical reports, or multi-stage reasoning traces, although applications should still request only the amount of output they need. Long responses increase latency and output cost, and a maximum limit is not the same as a recommendation to generate at that length.

Reasoning, coding, and tool use

Kimi K3 is aimed at reasoning-heavy work. The model supports reasoning-effort controls, allowing an application to manage the balance between deliberation and response efficiency. The supplied research describes it as suitable for long-horizon reasoning, autonomous coding, testing, iterative optimization, research programming, and extended agent workflows.

For coding, the model is intended to work beyond isolated code completion. A suitable workflow might provide a repository, ask Kimi K3 to identify relevant files, propose a change, write an implementation, run available tests through tools, inspect failures, and revise the result. The model’s long context is particularly relevant when the task spans multiple files or requires retaining project conventions and earlier debugging information.

Kimi K3 supports tool calling, so an application can expose functions such as search, file access, test execution, or external workflow actions. The model decides when to request a tool and can use the returned information in a subsequent response. Tool calling does not mean the hosted model has unrestricted access to a user’s systems: the surrounding application must define, authorize, and execute those tools.

The model also supports streaming responses, structured machine-readable responses, and context caching according to the supplied API information. Structured output can help software consume responses in a predictable format, while streaming allows an application to display partial text before the full response is complete. Caching may reduce the cost of repeated long prompts when the API and request pattern qualify.

Pricing and access options

The international Kimi API lists the following prices:

Usage typePrice per million tokens
Cache-miss input$3.00
Cache-hit input$0.30
Output$15.00

These are token-based API prices rather than a recurring consumer subscription price. A cache miss is input that must be processed normally; a cache hit refers to eligible repeated context served through the provider’s caching mechanism. Actual billing can depend on the endpoint, account, and region. A separate Chinese-language platform lists prices in yuan, so users should verify the price shown for their specific Kimi platform account before estimating costs.

Kimi K3 is also available through Kimi.com, Kimi Work, and Kimi Code, where access and quotas may be governed by the relevant product or membership arrangement rather than by the international API rates. Downloadable weights are available through Moonshot AI’s open-source release, but using them requires suitable inference infrastructure and operational expertise.

Main strengths and trade-offs

Kimi K3’s clearest strength is the combination of multimodal understanding and unusually large working context. It can bring together text, images, and video while retaining a very large amount of surrounding material. That combination is useful for codebase analysis, document-heavy research, visual debugging, and agents that need to work through several stages.

  • Long-context work: The one-million-token context window is suited to large repositories, extensive documentation, and long-running sessions.
  • Multimodal analysis: Native image and video input extends the model beyond text-only coding and research tasks.
  • Agentic coding: Tool calling, reasoning controls, and long output capacity support iterative implementation and testing workflows.
  • Open-weight availability: Organizations can evaluate downloadable weights when they need more control than a hosted-only model provides.
  • Structured integration: Streaming, structured responses, and caching are useful for production applications built around the API.

The main trade-off is resource demand. Kimi K3 is not the natural choice for inexpensive, low-latency chat or for deployment on ordinary local hardware. Its API output price is also much higher than its cache-miss input price, so applications that generate large responses should control output length and use caching where appropriate.

There are also operational limits around access and deployment. Advanced Kimi functions can depend on product, plan, plugin, or regional availability. The open-weight model’s total size can make self-hosting expensive and technically complex. In addition, the supplied research does not verify a model-specific knowledge cutoff, so current factual work may require external retrieval or tool use.

Best use cases for Kimi K3

Kimi K3 is a strong fit when the task combines large inputs, multiple reasoning steps, and a need for code or structured text output. Examples include:

  • Analyzing a large software repository and planning changes across many files.
  • Writing, testing, and iteratively debugging code through connected development tools.
  • Reviewing long technical documents, research materials, or collections of project files.
  • Combining screenshots, diagrams, video, and written requirements during investigation or design work.
  • Building research or engineering agents that search, call tools, maintain context, and produce a final report.
  • Generating detailed technical documentation or migration plans from extensive source material.

For these uses, the model’s value comes less from ordinary conversational fluency than from its ability to retain context and participate in a multi-step workflow.

When to choose Kimi K3

Choose Kimi K3 when long context, multimodal input, coding depth, or agentic tool use is more important than minimum latency and minimum cost. It is especially appropriate for teams that need to inspect large bodies of material in one working session or want to experiment with an open-weight model designed for complex reasoning.

A smaller or faster model may be more appropriate for routine classification, short customer-support replies, simple extraction, high-volume requests, or interactive applications where response speed matters more than extended reasoning. A text-only model may also be more economical when images and video are not part of the task. For native image, audio, or video generation, Kimi K3 is not the right choice because its output modality is text.

For hosted use, compare the API’s token costs with the expected input and output volume, paying particular attention to the $15.00 per million output-token price. For self-hosting, compare the control and customization benefits of downloadable weights with the cost of the required accelerator capacity and serving operations.

Kimi K3 specification summary

SpecificationReported detail
ProviderMoonshot AI
Model typeOpen-weight multimodal reasoning and coding model
Total parameters2.8 trillion
Activated parametersApproximately 104 billion
Context windowUp to 1,000,000 tokens
Maximum output131,072 tokens
InputsText, images, and video
OutputsText
Tool useSupported
StreamingSupported
Knowledge cutoffNot verified in the supplied official documentation

Overall, Kimi K3 is designed for demanding, context-heavy work rather than generic low-cost chat. Its combination of open-weight access, native visual understanding, long-context processing, coding support, and tool use makes it a noteworthy option for research and engineering workflows. The same design also brings higher infrastructure and usage costs, so its benefits are most apparent when a task genuinely requires that level of context and multi-step capability.


Answers to Frequently Asked Questions

How much does the Kimi K3 API cost?
The international Kimi API lists prices of $3.00 per million cache-miss input tokens, $0.30 per million cache-hit input tokens, and $15.00 per million output tokens. Actual pricing may vary by endpoint, account, region, or Kimi product, so users should verify the rates for their specific platform.
Is Kimi K3 suitable for coding and AI agents?
Yes. Kimi K3 supports reasoning controls, tool calling, streaming, structured responses, and context caching. It is intended for workflows such as inspecting repositories, writing and testing code, analyzing failures, revising implementations, and building research or engineering agents.
What inputs and outputs does Kimi K3 support?
Kimi K3 accepts text, images, and video as inputs and produces text as output. It can analyze screenshots, diagrams, visual documents, and video content, but it does not natively generate images, audio, or video.
What is Kimi K3?
Kimi K3 is an open-weight multimodal reasoning and coding model from Moonshot AI. It is designed for long-context tasks such as software engineering, technical research, document analysis, tool-assisted work, and autonomous or semi-autonomous agents.
How large is Kimi K3’s context window?
Kimi K3 supports a context window of up to 1,000,000 tokens. This makes it suitable for analyzing large codebases, lengthy technical documents, multiple files, and extended agent sessions, although practical performance depends on the task and serving infrastructure.


Sources 8
Provider

About Moonshot AI