Developer platform

Kimi API Platform

Developer overview for Moonshot AI, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API OpenAI-compatible Responses API
SDK support Official OpenAI Python and Node.js SDKs are supported through the OpenAI-compatible endpoints; official Anthropic SDKs are supported through the Anthropic-compatible Messages endpoint; raw HTTP works in any language.
Rate limits Limits are based on cumulative account recharge and account tier, with concurrency, RPM, TPM, TPD, and independent web-search QPS quotas. HTTP 429 responses include rate-limit headers. Exact quotas vary by tier and may be adjusted during capacity or risk-
Platform

API overview

Endpoints

API access

Base URL https://api.moonshot.ai/v1
Primary API OpenAI-compatible Responses API
Pricing

API pricing

Pricing model Pay-as-you-go token billing with separate pricing for model inference, web search, and batch processing

Input and output tokens are billed per usage; K3 supports cache-aware pricing, batch inference has separate lower-cost pricing, and file upload/extraction APIs are temporarily free.

Developer experience

SDKs & usability

SDKs Official OpenAI Python and Node.js SDKs are supported through the OpenAI-compatible endpoints; official Anthropic SDKs are supported through the Anthropic-compatible Messages endpoint; raw HTTP works in any language.
Ease of use Easy for developers familiar with OpenAI-compatible APIs; use an API key, change the base URL, select a current Kimi model, and send standard HTTP or SDK requests.
Documentation Good and currently maintained, with quickstarts, model guides, API references, examples, migration notices, pricing documentation, tool workflows, file workflows, and troubleshooting guidance.
Latency Model-dependent. Streaming reduces time to first token. The official model list describes kimi-k2.7-code-highspeed at approximately 180 tokens per second, with short-context peaks up to approximately 260 tokens per second; reasoning effort and model selec
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

The current platform supports OpenAI-compatible Chat Completions and Responses APIs, an Anthropic-compatible Messages API, streaming, multi-turn conversations, JSON Mode, structured output through response_format, function and tool calling, official web-search tools, dedicated web-search and URL-fetch endpoints, file upload and extraction, batch inference, context caching, and a browser-based Playground. Kimi K3, Kimi K2.7 Code, and Kimi K2.6 support multimodal input; supported models accept image input and the documentation also describes video input. The current model list should be used for production integrations. Moonshot V1, Kimi K2 preview models, Kimi K2.5, kimi-latest, and kimi-thinking-preview are listed as discontinued. The platform provides agent-building guidance using tool calling and official tools, but no separate persistent Assistants object is required. Fine-tuning was not verified in the current official API documentation.

Examples

API examples

curl --fail-with-body --silent --show-error https://api.moonshot.ai/v1/chat/completions \
  -H "Authorization: Bearer ${MOONSHOT_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain what an API is in one sentence."}
    ]
  }' | python -c 'import json,sys; print(json.load(sys.stdin)["choices"][0]["message"]["content"])'

curl --fail-with-body --silent --show-error https://api.moonshot.ai/v1/chat/completions \
  -H "Authorization: Bearer ${MOONSHOT_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "messages": [
      {"role": "system", "content": "Return only valid JSON."},
      {"role": "user", "content": "Return a JSON object with the keys name and type for the value Moonshot AI."}
    ],
    "response_format": {"type": "json_object"}
  }' | python -m json.tool
# Install or update the official compatible SDK:
# pip install --upgrade "openai>=1.0"

import json
import os
from openai import OpenAI

api_key = os.environ.get("MOONSHOT_API_KEY")
if not api_key:
    raise RuntimeError("Set MOONSHOT_API_KEY before running this example")

client = OpenAI(
    api_key=api_key,
    base_url="https://api.moonshot.ai/v1",
)

try:
    completion = client.chat.completions.create(
        model="kimi-k3",
        messages=[
            {"role": "system", "content": "You are a concise technical assistant."},
            {"role": "user", "content": "What is an API?"},
        ],
    )
    print(completion.choices[0].message.content)
except Exception as exc:
    print(f"Kimi API request failed: {exc}")

# Streaming example
try:
    stream = client.chat.completions.create(
        model="kimi-k3",
        messages=[
            {"role": "user", "content": "Explain event-driven architecture."},
        ],
        stream=True,
    )
    for chunk in stream:
        delta = chunk.choices[0].delta.content if chunk.choices else None
        if delta:
            print(delta, end="", flush=True)
    print()
except Exception as exc:
    print(f"Streaming request failed: {exc}")

# Structured JSON and function-tool example
weather_tool = {
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
            "additionalProperties": False,
        },
    },
}

response = client.chat.completions.create(
    model="kimi-k3",
    messages=[
        {"role": "system", "content": "Use tools when needed and return concise answers."},
        {"role": "user", "content": "What is the weather in Beijing?"},
    ],
    tools=[weather_tool],
    tool_choice="auto",
)

message = response.choices[0].message
if message.tool_calls:
    for tool_call in message.tool_calls:
        if tool_call.function.name == "get_weather":
            arguments = json.loads(tool_call.function.arguments)
            print(f"Tool requested for: {arguments['city']}")
else:
    print(message.content)

structured = client.chat.completions.create(
    model="kimi-k3",
    messages=[
        {"role": "system", "content": "Return only valid JSON."},
        {"role": "user", "content": "Return name and category for Moonshot AI."},
    ],
    response_format={"type": "json_object"},
)
print(json.loads(structured.choices[0].message.content))
// Install the official compatible SDK:
// npm install openai@latest

const OpenAI = require("openai");

const apiKey = process.env.MOONSHOT_API_KEY;
if (!apiKey) {
  throw new Error("Set MOONSHOT_API_KEY before running this example");
}

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.moonshot.ai/v1",
});

async function main() {
  try {
    const completion = await client.chat.completions.create({
      model: "kimi-k3",
      messages: [
        { role: "system", content: "You are a concise technical assistant." },
        { role: "user", content: "What is an API?" },
      ],
    });
    console.log(completion.choices[0].message.content);

    const structured = await client.chat.completions.create({
      model: "kimi-k3",
      messages: [
        { role: "system", content: "Return only valid JSON." },
        { role: "user", content: "Return name and category for Moonshot AI." },
      ],
      response_format: { type: "json_object" },
    });
    console.log(JSON.parse(structured.choices[0].message.content));

    const stream = await client.chat.completions.create({
      model: "kimi-k3",
      messages: [
        { role: "user", content: "Explain event-driven architecture." },
      ],
      stream: true,
    });
    for await (const chunk of stream) {
      const text = chunk.choices[0]?.delta?.content;
      if (text) process.stdout.write(text);
    }
    process.stdout.write("\n");
  } catch (error) {
    console.error("Kimi API request failed:", error.message);
    process.exitCode = 1;
  }
}

main();
<?php
$apiKey = getenv('MOONSHOT_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('Set MOONSHOT_API_KEY before running this example');
}

$url = 'https://api.moonshot.ai/v1/responses';
$payload = [
    'model' => 'kimi-k3',
    'input' => [
        [
            'role' => 'system',
            'content' => 'You are a concise technical assistant.'
        ],
        [
            'role' => 'user',
            'content' => 'Return a JSON object with the keys name and type for Moonshot AI.'
        ]
    ],
    'text' => [
        'format' => [
            'type' => 'json_object'
        ]
    ]
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_POST => true,
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json'
    ],
    CURLOPT_TIMEOUT => 60
]);

$raw = curl_exec($ch);
$curlError = curl_error($ch);
$httpCode = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

if ($raw === false || $curlError !== '') {
    throw new RuntimeException('Network error: ' . $curlError);
}
if ($httpCode < 200 || $httpCode >= 300) {
    throw new RuntimeException('Kimi API HTTP ' . $httpCode . ': ' . $raw);
}

$response = json_decode($raw, true, 512, JSON_THROW_ON_ERROR);

$text = $response['output_text'] ?? null;
if ($text === null && isset($response['output'])) {
    foreach ($response['output'] as $item) {
        foreach ($item['content'] ?? [] as $content) {
            if (isset($content['text'])) {
                $text = $content['text'];
                break 2;
            }
        }
    }
}

if ($text === null) {
    throw new RuntimeException('No text result found in the API response');
}

echo $text . PHP_EOL;
?>
Policies

Data & usage

Data training

The public Kimi OpenPlatform privacy policy states that Moonshot AI may use user content and related information to provide, improve, and develop its services, including training and refining underlying technology. A platform-specific no-training guarantee for ordinary API traffic was not verified. The documentation lists Zero Data Retention as an enterprise data-protection mode; eligibility and contractual details should be confirmed directly with Moonshot AI.

Data retention

The public privacy policy says information is retained as long as necessary for service provision, service improvement, dispute resolution, safety, security, and legal obligations, with retention varying by data type and account status. It states that account, input, and payment information may be retained while an account is active and that data may be stored on servers in Singapore. The platform documentation lists Zero Data Retention for enterprise customers, but the exact retention controls and eligibility should be verified with Moonshot AI.

Rate limits

Limits are based on cumulative account recharge and account tier, with concurrency, RPM, TPM, TPD, and independent web-search QPS quotas. HTTP 429 responses include rate-limit headers. Exact quotas vary by tier and may be adjusted during capacity or risk-

Developer guide

Moonshot AI API Guide: Getting Started with the Kimi API Platform

Moonshot AI's Kimi API Platform gives developers OpenAI-compatible Chat Completions and Responses APIs, an Anthropic-compatible Messages API, current Kimi models, streaming, structured JSON, tool calling, web search, file processing, batch inference, and a browser Playground. This beginner-friendly guide explains access, authentication, model selection, first requests, pricing, capabilities, SDK usage, limits, and production considerations.
The Kimi API Platform is Moonshot AI's developer service for adding Kimi models to applications. It is designed for teams building chat interfaces, coding tools, document workflows, research assistants, multimodal applications, and tool-using agents. The platform uses API keys and pay-as-you-go token billing, and its OpenAI-compatible endpoints make migration relatively straightforward for developers already familiar with standard chat APIs.
The Kimi API Platform provides OpenAI- and Anthropic-compatible interfaces for current Kimi models. Developers can use token-based billing, streaming, structured JSON, tool calling, multimodal input, file processing, web search, batch inference, SDKs, and the Kimi Playground while managing account-tier limits, data-use terms, and application-side state.

What is the Kimi API Platform?

The Kimi API Platform, also called Kimi Open Platform, lets software applications send requests to Moonshot AI's Kimi models and receive generated responses. An API, or application programming interface, is a structured way for one program to request work from another service. Instead of using Kimi through a browser or mobile app, developers use HTTP requests or an SDK inside their own applications.

The primary service URL is https://api.moonshot.ai. The platform provides OpenAI-compatible Chat Completions and Responses APIs, an Anthropic-compatible Messages API, file endpoints, web-search tools, URL fetching, and batch inference. For a new OpenAI-style integration, the current documentation recommends using a current Kimi model such as kimi-k3 through the OpenAI-compatible API.

The API is suitable for developers who need long-context text generation, coding assistance, document analysis, visual understanding, research workflows, structured responses, or applications that can call external tools. It is less suitable when an organization requires a fully documented no-training guarantee for ordinary API traffic, highly predictable quotas without account-tier requirements, or regional data-handling terms that have not been confirmed contractually.

Getting access and creating an API key

API access is managed through the Kimi Open Platform console. You need an account, an API key, and an account configuration that permits API use. The platform requires a small account recharge before use, so access is not equivalent to an unrestricted free developer tier.

Keep the key on a server or in a secret manager. Do not place it in browser JavaScript, a mobile application, source control, or client-side HTML where users can extract it. The examples below expect the key to be stored in an environment variable named MOONSHOT_API_KEY.

export MOONSHOT_API_KEY="your-api-key"

Every request uses a Bearer token in the Authorization header. If a request fails, first check that the key is present, the selected model is current, the account has access and balance, and the request is being sent to the correct endpoint.

Choosing an API and model

There are three main conversational interfaces:

  • Chat Completions: a familiar message-based API for standard conversations, streaming, vision input, JSON mode, and tool calls.
  • Responses API: a unified response interface for text or image inputs, structured output, function tools, and server-side web search.
  • Messages API: an Anthropic-compatible interface at /anthropic/v1/messages for applications using Anthropic-style request formats.

Use Chat Completions when you want the simplest OpenAI-compatible migration path. Use Responses when your application benefits from its unified response format, structured output, function tools, or hosted web search. Use the Messages API when compatibility with an Anthropic-oriented codebase is the main requirement.

The current model list includes kimi-k3, kimi-k2.7-code, kimi-k2.7-code-highspeed, and kimi-k2.6. Kimi K3 is the recommended general starting point for new applications and provides a 1-million-token context window with native visual understanding. Kimi K2.7 Code is aimed at software development and has a 256K context window; its high-speed variant is intended for latency-sensitive coding workloads. Kimi K2.6 supports text, image, and video input, with thinking and non-thinking modes.

Check the current model documentation before deploying a model identifier. The supplied current model guidance lists Moonshot V1, Kimi K2.5, Kimi K2 preview models, kimi-latest, and kimi-thinking-preview as discontinued or unsuitable for new integrations.

Make your first request

The following request uses the OpenAI-compatible Chat Completions endpoint. The API version is selected by changing the base URL to https://api.moonshot.ai/v1.

curl --fail-with-body --silent --show-error https://api.moonshot.ai/v1/chat/completions 
  -H "Authorization: Bearer ${MOONSHOT_API_KEY}" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "kimi-k3",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain what an API is in one sentence."}
    ]
  }'

A successful Chat Completions response contains a list of choices. The generated text is normally read from choices[0].message.content. The response also includes metadata such as the model and usage information, which you can use to monitor token consumption.

Use the OpenAI-compatible Python SDK

Moonshot AI supports the official OpenAI Python SDK when you provide the Kimi base URL. The SDK package and client syntax below use the current OpenAI 1.x-style interface.

import os
from openai import OpenAI

api_key = os.environ.get("MOONSHOT_API_KEY")
if not api_key:
    raise RuntimeError("Set MOONSHOT_API_KEY before running this example")

client = OpenAI(
    api_key=api_key,
    base_url="https://api.moonshot.ai/v1",
)

response = client.chat.completions.create(
    model="kimi-k3",
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "What is an API?"},
    ],
)

print(response.choices[0].message.content)

Official OpenAI Node.js SDKs can be configured in the same general way by setting the Kimi base URL. Official Anthropic SDKs are supported through the Anthropic-compatible Messages endpoint. Raw HTTP requests are also available from any programming language, including PHP.

How Kimi API pricing works

Kimi API uses pay-as-you-go billing rather than a normal monthly subscription for inference. Model usage is billed according to input and output token consumption. Input tokens represent the content sent to the model, while output tokens represent generated content. Long prompts, conversation history, retrieved documents, and extracted files can therefore increase input usage.

K3 models support cache-aware pricing, with cache-write and cached-input behavior connected to cache time-to-live settings. Batch inference has separate pricing and is intended for asynchronous bulk processing at a lower cost than ordinary requests. File upload and extraction APIs are temporarily free according to the supplied documentation, but text extracted from files and passed into a model is billed as input.

Exact token prices are model-specific and can change, so consult the current pricing documentation rather than embedding a price in application logic or relying on an old example. Web-search requests also have separate pricing and quota considerations.

Core capabilities available to developers

Streaming responses

Streaming sends generated content incrementally over server-sent events instead of waiting for the complete answer. It is useful for chat interfaces because the application can display the beginning of a response while the model continues generating.

stream = client.chat.completions.create(
    model="kimi-k3",
    messages=[
        {"role": "user", "content": "Explain event-driven architecture."},
    ],
    stream=True,
)

for chunk in stream:
    if chunk.choices:
        text = chunk.choices[0].delta.content
        if text:
            print(text, end="", flush=True)
print()

Structured JSON output

JSON Mode asks the model to return a JSON object instead of ordinary prose. In Chat Completions, set response_format to {"type": "json_object"} and still validate and parse the result in application code. JSON Mode does not remove the need for error handling: the application should handle malformed output, missing fields, unexpected values, and service errors.

import json

response = client.chat.completions.create(
    model="kimi-k3",
    messages=[
        {"role": "system", "content": "Return only valid JSON."},
        {"role": "user", "content": "Return name and category for Moonshot AI."},
    ],
    response_format={"type": "json_object"},
)

result = json.loads(response.choices[0].message.content)
print(result)

The Responses API also documents structured output through its response-format interface. Use the interface supported by the endpoint and model you select, and validate the decoded object against your own application schema.

Function and tool calling

Function calling lets the model request an operation that your application performs. For example, the model can request a weather lookup, database query, or internal business action. Your application remains responsible for executing the function, checking permissions, validating arguments, and returning the result. The model does not automatically gain access to your systems merely because a function is declared.

weather_tool = {
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
            "additionalProperties": False,
        },
    },
}

response = client.chat.completions.create(
    model="kimi-k3",
    messages=[
        {"role": "user", "content": "What is the weather in Beijing?"},
    ],
    tools=[weather_tool],
    tool_choice="auto",
)

message = response.choices[0].message
if message.tool_calls:
    for tool_call in message.tool_calls:
        if tool_call.function.name == "get_weather":
            arguments = json.loads(tool_call.function.arguments)
            print("Requested city:", arguments["city"])
else:
    print(message.content)

A production tool loop normally sends the tool result back in a follow-up request so the model can produce a final answer. Preserve tool-call identifiers, validate JSON arguments, restrict available operations, and avoid allowing model-generated arguments to bypass authorization.

Images, video, and files

Supported Kimi models accept multimodal input. Image data can be provided as a URL or base64 data URL. The current documentation also describes video input for supported models. The platform provides file upload and extraction workflows for documents, images, and video. Extracted document text is included in the model request and is billed as input tokens even when the upload and extraction operation itself is temporarily free.

File processing is useful for document question answering, research, spreadsheet and report workflows, and visual analysis. Applications should still control file size, file type, sensitive information, retention expectations, and the amount of extracted content included in each request.

Web search and URL fetching

The platform offers dedicated web-search and URL-fetch endpoints as well as official web-search tools for supported model workflows. These features can support research and current-information applications, but search requests have separate pricing and rate-limit considerations. Treat retrieved pages as untrusted external content and apply the same validation and prompt-injection protections used for other web-connected systems.

Playground and developer tools

The browser-based Kimi Playground lets developers test prompts, compare current models, tune parameters, try tools, inspect token usage, and generate request code. It is useful for exploring a prompt or API feature before writing a complete application.

The official documentation includes quickstarts, API references, model information, pricing guidance, migration notices, file examples, tool-calling instructions, error references, and troubleshooting material. Use the current model list and documentation when moving from Playground experiments to production because older examples may reference discontinued model families.

Limits and production considerations

Rate limits depend on account recharge and account tier. They can include concurrency, requests per minute, tokens per minute, and tokens per day. Web-search endpoints use an independent queries-per-second quota. Exact quotas can vary by account and may change during capacity or risk-control events.

When a quota is exceeded, the API can return HTTP 429 along with rate-limit headers. Production clients should inspect those headers, retry only when appropriate, and use exponential backoff with jitter. Also set request timeouts, log request identifiers and usage safely, and monitor latency, token consumption, cache behavior, model errors, and failed tool calls.

The platform's core conversation API is stateless. If an application needs conversation history, user preferences, permissions, or durable agent state, the application must store and manage that information. Agent-building guidance does not require a separate persistent Assistants object; developers generally implement persistence, tool execution, authorization, and retry behavior themselves.

Privacy terms require particular attention. The public Kimi OpenPlatform privacy policy says Moonshot AI may use prompts, files, generated content, account information, and usage data to provide, improve, and develop its services, including training and refining underlying technology. A general no-training guarantee for ordinary API traffic was not verified in the supplied documentation. Zero Data Retention is listed as an enterprise data-protection mode, but eligibility, contractual terms, retention controls, and regional processing arrangements should be confirmed directly with Moonshot AI before sending sensitive information.

When Kimi API is a good or poor choice

Kimi API is a strong candidate when your team wants an OpenAI-compatible migration path, very long context for documents or code, multimodal input, hosted web-search options, tool calling, JSON responses, file workflows, or access to both standard and coding-focused Kimi models. Its compatibility with OpenAI and Anthropic SDK patterns can reduce the amount of integration code required.

It may be a poor fit when your application needs fixed, universally available quotas, a clearly verified no-training policy for ordinary API traffic, guaranteed regional processing outside the documented arrangements, or a fully managed persistent agent state. Teams with strict compliance requirements should review the privacy policy and enterprise terms before production use. Teams with predictable high-volume workloads should also compare current model, batch, search, and cache pricing and confirm account-tier limits.

For a first project, create a protected API key, start with kimi-k3, make a simple Chat Completions request, inspect the response and usage fields, then add streaming, structured output, tools, files, or web search one feature at a time. Recheck the official model list and migration notices before deployment.

Sources 13