What is the Kimi API Platform?
The Kimi API Platform, also called Kimi Open Platform, lets software applications send requests to Moonshot AI's Kimi models and receive generated responses. An API, or application programming interface, is a structured way for one program to request work from another service. Instead of using Kimi through a browser or mobile app, developers use HTTP requests or an SDK inside their own applications.
The primary service URL is https://api.moonshot.ai. The platform provides OpenAI-compatible Chat Completions and Responses APIs, an Anthropic-compatible Messages API, file endpoints, web-search tools, URL fetching, and batch inference. For a new OpenAI-style integration, the current documentation recommends using a current Kimi model such as kimi-k3 through the OpenAI-compatible API.
The API is suitable for developers who need long-context text generation, coding assistance, document analysis, visual understanding, research workflows, structured responses, or applications that can call external tools. It is less suitable when an organization requires a fully documented no-training guarantee for ordinary API traffic, highly predictable quotas without account-tier requirements, or regional data-handling terms that have not been confirmed contractually.
Getting access and creating an API key
API access is managed through the Kimi Open Platform console. You need an account, an API key, and an account configuration that permits API use. The platform requires a small account recharge before use, so access is not equivalent to an unrestricted free developer tier.
Keep the key on a server or in a secret manager. Do not place it in browser JavaScript, a mobile application, source control, or client-side HTML where users can extract it. The examples below expect the key to be stored in an environment variable named MOONSHOT_API_KEY.
export MOONSHOT_API_KEY="your-api-key"Every request uses a Bearer token in the Authorization header. If a request fails, first check that the key is present, the selected model is current, the account has access and balance, and the request is being sent to the correct endpoint.
Choosing an API and model
There are three main conversational interfaces:
- Chat Completions: a familiar message-based API for standard conversations, streaming, vision input, JSON mode, and tool calls.
- Responses API: a unified response interface for text or image inputs, structured output, function tools, and server-side web search.
- Messages API: an Anthropic-compatible interface at
/anthropic/v1/messagesfor applications using Anthropic-style request formats.
Use Chat Completions when you want the simplest OpenAI-compatible migration path. Use Responses when your application benefits from its unified response format, structured output, function tools, or hosted web search. Use the Messages API when compatibility with an Anthropic-oriented codebase is the main requirement.
The current model list includes kimi-k3, kimi-k2.7-code, kimi-k2.7-code-highspeed, and kimi-k2.6. Kimi K3 is the recommended general starting point for new applications and provides a 1-million-token context window with native visual understanding. Kimi K2.7 Code is aimed at software development and has a 256K context window; its high-speed variant is intended for latency-sensitive coding workloads. Kimi K2.6 supports text, image, and video input, with thinking and non-thinking modes.
Check the current model documentation before deploying a model identifier. The supplied current model guidance lists Moonshot V1, Kimi K2.5, Kimi K2 preview models, kimi-latest, and kimi-thinking-preview as discontinued or unsuitable for new integrations.
Make your first request
The following request uses the OpenAI-compatible Chat Completions endpoint. The API version is selected by changing the base URL to https://api.moonshot.ai/v1.
curl --fail-with-body --silent --show-error https://api.moonshot.ai/v1/chat/completions
-H "Authorization: Bearer ${MOONSHOT_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "kimi-k3",
"messages": [
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain what an API is in one sentence."}
]
}'A successful Chat Completions response contains a list of choices. The generated text is normally read from choices[0].message.content. The response also includes metadata such as the model and usage information, which you can use to monitor token consumption.
Use the OpenAI-compatible Python SDK
Moonshot AI supports the official OpenAI Python SDK when you provide the Kimi base URL. The SDK package and client syntax below use the current OpenAI 1.x-style interface.
import os
from openai import OpenAI
api_key = os.environ.get("MOONSHOT_API_KEY")
if not api_key:
raise RuntimeError("Set MOONSHOT_API_KEY before running this example")
client = OpenAI(
api_key=api_key,
base_url="https://api.moonshot.ai/v1",
)
response = client.chat.completions.create(
model="kimi-k3",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "What is an API?"},
],
)
print(response.choices[0].message.content)Official OpenAI Node.js SDKs can be configured in the same general way by setting the Kimi base URL. Official Anthropic SDKs are supported through the Anthropic-compatible Messages endpoint. Raw HTTP requests are also available from any programming language, including PHP.
How Kimi API pricing works
Kimi API uses pay-as-you-go billing rather than a normal monthly subscription for inference. Model usage is billed according to input and output token consumption. Input tokens represent the content sent to the model, while output tokens represent generated content. Long prompts, conversation history, retrieved documents, and extracted files can therefore increase input usage.
K3 models support cache-aware pricing, with cache-write and cached-input behavior connected to cache time-to-live settings. Batch inference has separate pricing and is intended for asynchronous bulk processing at a lower cost than ordinary requests. File upload and extraction APIs are temporarily free according to the supplied documentation, but text extracted from files and passed into a model is billed as input.
Exact token prices are model-specific and can change, so consult the current pricing documentation rather than embedding a price in application logic or relying on an old example. Web-search requests also have separate pricing and quota considerations.
Core capabilities available to developers
Streaming responses
Streaming sends generated content incrementally over server-sent events instead of waiting for the complete answer. It is useful for chat interfaces because the application can display the beginning of a response while the model continues generating.
stream = client.chat.completions.create(
model="kimi-k3",
messages=[
{"role": "user", "content": "Explain event-driven architecture."},
],
stream=True,
)
for chunk in stream:
if chunk.choices:
text = chunk.choices[0].delta.content
if text:
print(text, end="", flush=True)
print()Structured JSON output
JSON Mode asks the model to return a JSON object instead of ordinary prose. In Chat Completions, set response_format to {"type": "json_object"} and still validate and parse the result in application code. JSON Mode does not remove the need for error handling: the application should handle malformed output, missing fields, unexpected values, and service errors.
import json
response = client.chat.completions.create(
model="kimi-k3",
messages=[
{"role": "system", "content": "Return only valid JSON."},
{"role": "user", "content": "Return name and category for Moonshot AI."},
],
response_format={"type": "json_object"},
)
result = json.loads(response.choices[0].message.content)
print(result)The Responses API also documents structured output through its response-format interface. Use the interface supported by the endpoint and model you select, and validate the decoded object against your own application schema.
Function and tool calling
Function calling lets the model request an operation that your application performs. For example, the model can request a weather lookup, database query, or internal business action. Your application remains responsible for executing the function, checking permissions, validating arguments, and returning the result. The model does not automatically gain access to your systems merely because a function is declared.
weather_tool = {
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
"additionalProperties": False,
},
},
}
response = client.chat.completions.create(
model="kimi-k3",
messages=[
{"role": "user", "content": "What is the weather in Beijing?"},
],
tools=[weather_tool],
tool_choice="auto",
)
message = response.choices[0].message
if message.tool_calls:
for tool_call in message.tool_calls:
if tool_call.function.name == "get_weather":
arguments = json.loads(tool_call.function.arguments)
print("Requested city:", arguments["city"])
else:
print(message.content)A production tool loop normally sends the tool result back in a follow-up request so the model can produce a final answer. Preserve tool-call identifiers, validate JSON arguments, restrict available operations, and avoid allowing model-generated arguments to bypass authorization.
Images, video, and files
Supported Kimi models accept multimodal input. Image data can be provided as a URL or base64 data URL. The current documentation also describes video input for supported models. The platform provides file upload and extraction workflows for documents, images, and video. Extracted document text is included in the model request and is billed as input tokens even when the upload and extraction operation itself is temporarily free.
File processing is useful for document question answering, research, spreadsheet and report workflows, and visual analysis. Applications should still control file size, file type, sensitive information, retention expectations, and the amount of extracted content included in each request.
Web search and URL fetching
The platform offers dedicated web-search and URL-fetch endpoints as well as official web-search tools for supported model workflows. These features can support research and current-information applications, but search requests have separate pricing and rate-limit considerations. Treat retrieved pages as untrusted external content and apply the same validation and prompt-injection protections used for other web-connected systems.
Playground and developer tools
The browser-based Kimi Playground lets developers test prompts, compare current models, tune parameters, try tools, inspect token usage, and generate request code. It is useful for exploring a prompt or API feature before writing a complete application.
The official documentation includes quickstarts, API references, model information, pricing guidance, migration notices, file examples, tool-calling instructions, error references, and troubleshooting material. Use the current model list and documentation when moving from Playground experiments to production because older examples may reference discontinued model families.
Limits and production considerations
Rate limits depend on account recharge and account tier. They can include concurrency, requests per minute, tokens per minute, and tokens per day. Web-search endpoints use an independent queries-per-second quota. Exact quotas can vary by account and may change during capacity or risk-control events.
When a quota is exceeded, the API can return HTTP 429 along with rate-limit headers. Production clients should inspect those headers, retry only when appropriate, and use exponential backoff with jitter. Also set request timeouts, log request identifiers and usage safely, and monitor latency, token consumption, cache behavior, model errors, and failed tool calls.
The platform's core conversation API is stateless. If an application needs conversation history, user preferences, permissions, or durable agent state, the application must store and manage that information. Agent-building guidance does not require a separate persistent Assistants object; developers generally implement persistence, tool execution, authorization, and retry behavior themselves.
Privacy terms require particular attention. The public Kimi OpenPlatform privacy policy says Moonshot AI may use prompts, files, generated content, account information, and usage data to provide, improve, and develop its services, including training and refining underlying technology. A general no-training guarantee for ordinary API traffic was not verified in the supplied documentation. Zero Data Retention is listed as an enterprise data-protection mode, but eligibility, contractual terms, retention controls, and regional processing arrangements should be confirmed directly with Moonshot AI before sending sensitive information.
When Kimi API is a good or poor choice
Kimi API is a strong candidate when your team wants an OpenAI-compatible migration path, very long context for documents or code, multimodal input, hosted web-search options, tool calling, JSON responses, file workflows, or access to both standard and coding-focused Kimi models. Its compatibility with OpenAI and Anthropic SDK patterns can reduce the amount of integration code required.
It may be a poor fit when your application needs fixed, universally available quotas, a clearly verified no-training policy for ordinary API traffic, guaranteed regional processing outside the documented arrangements, or a fully managed persistent agent state. Teams with strict compliance requirements should review the privacy policy and enterprise terms before production use. Teams with predictable high-volume workloads should also compare current model, batch, search, and cache pricing and confirm account-tier limits.
For a first project, create a protected API key, start with kimi-k3, make a simple Chat Completions request, inspect the response and usage fields, then add streaming, structured output, tools, files, or web search one feature at a time. Recheck the official model list and migration notices before deployment.
