What is Kimi K2.8 Preview?
Kimi K2.8 Preview is Moonshot AI’s preview model for coding and software-development agents. It was fully rolled out in Kimi Code on September 11, 2026, replacing the earlier model behind the existing kimi-for-coding serving identifier. The display name is Kimi K2.8 Preview, but the identifier used in supported configurations is kimi-for-coding.
This distinction matters in practice. Users of Kimi Code, its command-line interface, the VS Code extension, the desktop application, or compatible third-party coding tools generally do not need to replace an existing kimi-for-coding configuration. Entering Kimi K2.8 Preview as though it were the serving ID may cause a request to fail.
Moonshot positions K2.8 Preview close to Kimi K3 for coding and agent tasks, while describing it as more efficient in its thinking process than K2.7 Code. Those are provider positioning claims rather than independently verified benchmark results. No public benchmark table or parameter count for this exact preview model was identified in the supplied documentation.
A 1-million-token context window and adjustable thinking
The model supports a context window of up to 1,048,576 tokens. A context window is the amount of text and other request information the model can consider at once. In practical terms, the documented limit is intended to support large repositories, long technical documents, extensive issue histories, and multi-file coding sessions without requiring the user to summarize as aggressively.
Kimi Code documentation states that the 1-million-token context limit is available across membership tiers that include access to the model, although the exact model and context entitlements remain dependent on the account plan. A large context window does not guarantee that every request will be processed identically: quotas, account limits, tool activity, and the size of the requested response can still affect a session.
K2.8 Preview provides three documented thinking-effort levels: low, high, and max. The default is max. Thinking can also be disabled. These controls let users choose between deeper reasoning and lower latency, although Moonshot has not supplied a fixed response-time or quality guarantee for each setting.
When thinking is disabled for K3-series or K2.8 Preview requests, Kimi Code routes the request to a non-thinking K2.8 Preview configuration. This makes the model more suitable for routine transformations, straightforward code completion, or interactive tasks where waiting for extended reasoning is less useful.
Coding and agent capabilities
K2.8 Preview primarily produces text: source code, explanations, patches, plans, error analysis, and agent responses. Its intended workload is broader than single-function code generation. It is suited to examining a repository, tracing how files relate to one another, proposing a change, and using tools as part of an extended development task.
- Repository analysis: the large context window can accommodate substantial codebases or long collections of project files, subject to the applicable account and request limits.
- Code completion and generation: the model can produce implementation code, tests, explanations, and suggested changes.
- Multi-file refactoring: it can reason across related files and describe or apply coordinated edits through a coding agent.
- Long-horizon tasks: the model is positioned for tasks involving several investigative or implementation steps rather than one isolated answer.
- Tool-assisted development: Kimi Code exposes tool use for coding-agent workflows, allowing the surrounding client to manage actions and return results to the model.
The available documentation confirms tool use and streaming. Streaming means that a client can receive parts of the response as they are generated rather than waiting for the complete response. This is useful in interactive coding tools, but it does not establish a specific latency target.
Supported input and output types
The model’s native output is text. It does not generate images, audio, or video as model output. Kimi Code configuration documentation identifies image and video input capabilities, however, making the model usable for some multimodal coding tasks.
For example, a coding agent may be able to inspect a screenshot of a user-interface problem, a diagram, or a video-related development asset alongside source code. The supplied research does not document audio input for K2.8 Preview, so audio should not be assumed to be supported.
| Capability | Documented status |
|---|---|
| Text input | Supported |
| Image input | Supported through Kimi Code documentation |
| Video input | Supported through Kimi Code documentation |
| Text output | Supported |
| Image, audio, or video output | Not supported as native model output |
| Tool use | Supported |
| Streaming | Supported |
These modality details describe the model and its Kimi Code integration, not a claim that every client exposes every input type in the same way.
How to access Kimi K2.8 Preview
Kimi K2.8 Preview is available through Kimi Code clients, including the desktop application, CLI, VS Code extension, and supported third-party coding tools. Kimi Code provides OpenAI-compatible endpoints at https://api.kimi.com/coding/v1 and https://api.kimi.ai/coding/v1, along with Anthropic-compatible endpoints.
The canonical model identifier remains kimi-for-coding. Kimi Code membership entitlements and quota rules control access and usage. Moonshot has not published a standalone per-token input or output price for K2.8 Preview in the supplied documentation, so there is no verified model-specific token price to quote.
This pricing structure makes the model different from a separately listed API model with a public input and output rate. Before adopting it for automated workloads, users should verify the relevant membership tier, quota, endpoint terms, and client support. The preview status and non-versioned serving identifier also mean that future routing changes may occur without a new identifier.
How it compares with other coding options
K2.8 Preview is aimed at users who value context capacity and deeper coding-agent reasoning more than the lowest possible latency. Its adjustable thinking settings provide a practical way to make that trade-off: max is better suited to difficult repository-level work, while thinking-disabled or lower-effort requests may be preferable for simple, repetitive tasks.
Moonshot describes K2.8 Preview as close to Kimi K3 in coding and agent capability. It is also distinct from K2.7 Code HighSpeed. HighSpeed is a separate model ID intended for substantially faster output, whereas K2.8 Preview is the standard coding model with the larger documented context window and adjustable thinking effort. The supplied research does not provide a controlled benchmark or price comparison between the two, so the choice should be based on the actual latency, quota, and task requirements of the account.
For a workflow that needs native image, audio, or video generation, K2.8 Preview is not the appropriate choice. It is also a weaker fit when a project requires a stable, separately versioned model name, a published per-token price, a documented maximum output-token limit, or a published knowledge cutoff. Those details were not provided for this preview model.
Strengths and limitations
Where K2.8 Preview is strongest
- Its up-to-1,048,576-token context window is well suited to large repositories, long specifications, and extended coding sessions.
- It combines coding focus with agent-oriented tool use rather than limiting the interaction to isolated code snippets.
- Adjustable thinking effort gives users control over the quality-versus-latency trade-off.
- Existing Kimi Code setups using
kimi-for-codingcan generally continue using the same identifier after the model replacement. - Image and video input can help with software tasks involving screenshots, diagrams, or related visual assets.
Important limitations
- It is a preview model, so behavior and routing may change.
- Access is tied to Kimi Code membership benefits and quotas rather than a separately published token price.
- No verified maximum output-token limit, parameter count, knowledge cutoff, fine-tuning support, caching support, or batch API was documented for this exact model.
- It produces text rather than native image, audio, or video output.
- Independent benchmark results were not available in the supplied sources; comparisons such as being close to Kimi K3 are provider claims.
- The model’s non-versioned
kimi-for-codingidentifier may route to a future model without a configuration change.
When to choose Kimi K2.8 Preview
Choose Kimi K2.8 Preview when the task involves substantial code context, repository-wide reasoning, multi-file changes, or an agent that needs to work through several development steps. It is particularly suitable when a Kimi Code membership already provides access and when the user can benefit from switching between lower-effort interactive responses and deeper reasoning for difficult problems.
A faster coding model such as K2.7 Code HighSpeed may be more appropriate when response speed is the primary requirement and the task does not need the same context capacity or reasoning depth. A separately priced or versioned API model may be preferable for production systems that need predictable per-token billing, stable model lifecycle controls, or documented output limits. A model designed for media generation should be used instead when the required output is an image, audio file, or video.
Overall, K2.8 Preview is best understood as a large-context coding-agent option within Kimi Code, not as a general-purpose media model or a conventional standalone API model. Its main practical advantages are repository-scale context, tool-assisted development, and controllable reasoning; its main uncertainties are preview status, quota-based access, and the absence of several published model specifications.

