Model catalog

Moonshot AI Models

Browse the AI models associated with Moonshot AI. Compare current and historical models by family, capabilities, context window, availability and intended use.

19 models tracked
19 Total models
9 Model families
6 Model types
18 Current / accessible
All models

Moonshot AI model catalog

Moonshot AI logo
Kimi-Audio

Kimi-Audio-7B

Fine-tuning and research on speech recognition, audio understanding, audio classification, audio question answering, and speech-audio generation

Type Multimodal
Context 8K
Reasoning 4/10
Speed 5/10
Multimodal Audio input Media output
Status

Current open-weight base model; not instruction-tuned

Input No official hosted API pricing; downloadable open weights
Output No official hosted API pricing; downloadable open weights
View model →

Self-hosted speech recognition, audio understanding, audio question answering, audio captioning, and spoken conversational agents

Type Multimodal
Reasoning 6/10
Speed 5/10
Multimodal Audio input Media output
Status

Available as an open-weight self-hosted model

Input No official hosted API pricing
Output No official hosted API pricing
View model →

Repository-level software issue resolution, code repair, test generation, coding agents, and self-hosted software engineering workflows

Type Coding
Context 131K
Reasoning 7/10
Speed 4/10
Status

Available open-weight model

View model →

Self-hosted image and video understanding, OCR, long-document analysis, visual question answering, screenshot perception, and efficient multimodal applications

Type Multimodal
Context 131K
Reasoning 7/10
Speed 8/10
Multimodal Image input Video input
Status

Current open-weight/downloadable model; a newer Kimi-VL-A3B-Thinking-2506 variant is recommended for stronger multimodal reasoning

Input No official Moonshot-hosted API price published; self-hosted/open-weight model
Output No official Moonshot-hosted API price published; self-hosted/open-weight model
View model →

Local multimodal reasoning, mathematical visual question answering, OCR, document understanding, image and video analysis, and research on open-weight vision-language models

Type Reasoning
Context 131K
Reasoning 7/10
Speed 7/10
Multimodal Image input Video input
Status

Available as downloadable open weights; superseded by Kimi-VL-A3B-Thinking-2506

View model →

Open-weight image and video reasoning, OCR, chart interpretation, visual mathematics, long PDFs, high-resolution screenshots and GUI-agent grounding

Type Reasoning
Context 131K
Reasoning 8/10
Speed 7/10
Multimodal Image input Video input
Status

Current open-weight model; publicly available for self-hosted and third-party inference

Input No official hosted API price; open-weight model
Output No official hosted API price; open-weight model
View model →

Fine-tuning, foundation-model research, custom language systems, coding experiments, and self-hosted inference

Type General Purpose
Context 131K
Reasoning 8/10
Speed 3/10
Fine-tuning Streaming
Status

Open-weight and downloadable; legacy relative to newer Kimi K2.x models but still accessible from the official Hugging Face repository

Input No official hosted API price for the Base checkpoint; self-hosted weights
Output No official hosted API price for the Base checkpoint; self-hosted weights
View model →

Open-weight coding assistants, tool-using agents, general-purpose chat, and self-hosted research deployments

Type General Purpose
Context 131K
Reasoning 8/10
Speed 6/10
Tool use Streaming
Status

Open-weight checkpoint available; legacy relative to Moonshot AI's current hosted API catalog

Input USD 0.60 per 1 million input tokens historically for the Kimi K2 preview API; current first-party pricing for this exact model is unverified
Output USD 2.50 per 1 million output tokens historically for the Kimi K2 preview API; current first-party pricing for this exact model is unverified
View model →

Self-hosted reasoning agents, autonomous research, long-horizon tool workflows, coding, and complex multi-step analysis

Type Reasoning
Context 262K
Reasoning 9/10
Speed 5/10
Tool use Streaming
Status

Retired from Moonshot's direct API; open-weight checkpoint remains available for self-hosted or third-party deployment

View model →
Moonshot AI logo
Kimi K2

Kimi K2.6

Long-horizon software engineering, agentic coding, visual document understanding, tool-using workflows, and multi-agent orchestration

Type Multimodal
Context 262K
Reasoning 8/10
Speed 8/10
Multimodal Image input Tool use
Status

Current; open-source model, available through Kimi, Kimi API, Kimi Code, and downloadable model weights

View model →

Long-horizon software engineering, repository-level coding, multi-file refactoring, debugging, coding agents, and tool-driven development workflows

Type Coding
Context 262K
Reasoning 8/10
Speed 7/10
Multimodal Image input Video input
Status

Current and accessible through Kimi API; Kimi Code default service has moved to Kimi K2.8 Preview, while Kimi K2.7 Code remains available through API and Kimi K2.7 Code HighSpeed service.

Input ¥6.50 per 1M uncached input tokens; ¥1.30 per 1M cache-hit input tokens
Output ¥27.00 per 1M output tokens
View model →

Long-context software development, repository analysis, code completion, multi-file refactoring, and agentic coding workflows

Type Coding
Context 1.05M
Reasoning 8/10
Speed 7/10
Multimodal Image input Video input
Status

Preview; fully rolled out in Kimi Code

Input Not publicly listed; accessed through Kimi Code membership and quota
Output Not publicly listed; accessed through Kimi Code membership and quota
View model →
Moonshot AI logo
Kimi K2.5

Kimi K2.5

Multimodal coding, visual debugging, long-context analysis, tool-using agents, and complex research or office workflows

Type Multimodal
Context 256K
Reasoning 9/10
Speed 8/10
Multimodal Image input Video input
Status

current

Input $0.60 per 1 million tokens; cached input approximately $0.10 per 1 million tokens
Output $3.00 per 1 million tokens
View model →
Moonshot AI logo
Kimi K3

Kimi K3

Long-context coding, software engineering, multimodal document and video understanding, agentic workflows, technical research, and complex reasoning

Type Multimodal
Context 1M
Reasoning 9/10
Speed 5/10
Multimodal Image input Video input
Status

Current and available; open-weight model

Input $3.00 per MTok cache-miss input; $0.30 per MTok cache-hit input
Output $15.00 per MTok
View model →

Long-context text generation, local inference, continued pretraining, research, and fine-tuning

Type Lightweight
Context 1.05M
Reasoning 7/10
Speed 9/10
Fine-tuning Streaming
Status

Current open-weight model; publicly downloadable

Input Not applicable; open-weight checkpoint with no official hosted API price
Output Not applicable; open-weight checkpoint with no official hosted API price
View model →

Long-context text generation, self-hosted assistants, document analysis, research, and efficient local inference

Type General Purpose
Context 1.05M
Reasoning 7/10
Speed 9/10
Streaming
Status

Current open-weight instruction-tuned checkpoint

View model →

Self-hosted text generation, language-model research, efficient MoE inference, code and mathematics experimentation, and fine-tuning research

Type General Purpose
Context 8K
Reasoning 6/10
Speed 8/10
Status

Available open-weight pretrained checkpoint

View model →

Self-hosted instruction following, general text generation, research, experimentation, and cost-conscious local inference

Type General Purpose
Context 8K
Reasoning 6/10
Speed 8/10
Status

Available open-weight checkpoint

View model →

Native-resolution image feature extraction, vision-language model backbones, high-resolution document and image understanding, and multimodal research

Type Other
Image input
Status

Current open-weight model; available for local use through Hugging Face Transformers; not deployed by a Hugging Face Inference Provider

View model →