What is the Xiaomi MiMo API?
The Xiaomi MiMo API Open Platform is Xiaomi's current developer-facing model-inference platform. It lets applications send prompts and other supported inputs to MiMo models and receive generated responses over HTTPS.
The platform is designed to be familiar to developers who have used other model APIs. Its primary interface is OpenAI-compatible, and Xiaomi also documents Anthropic-compatible access. The main real-time base URL is https://api.xiaomimimo.com/v1.
Use the API when you want to add model-generated text, image understanding, tool use, web-assisted answers, structured JSON, or asynchronous batch processing to an application. It is aimed at developers building services, automation, agent workflows, coding tools, and other software integrations rather than people looking for a standalone consumer chatbot.
How to get access and create an API key
Start by creating or using a Xiaomi account and opening the MiMo API Open Platform console. From the console, create an API key and check which billing mode and regional service are associated with it.
Pay-as-you-go API keys use the sk- format. Token Plan credentials use separate tp- or team-plan formats and must be used with their dedicated base URL. These credential types and balances are not interchangeable.
Send the key with either the api-key header or a Bearer authorization header, depending on the integration pattern. Store it in an environment variable or secret manager rather than placing it in browser code, a source repository, or a mobile application.
Xiaomi's current documentation primarily describes personal-account login, and some services may require additional account or real-name verification. Check the console before planning an automated production signup process.
Which model and API should you choose?
For new integrations, start with the current MiMo V2.6 family:
mimo-v2.6-prois the general higher-capability option and supports full-modality use cases.mimo-v2.6-flashis the lower-cost option for many text and multimodal workloads.mimo-v2.6-pro-ultraspeedis intended for latency-sensitive workloads, but its availability and service terms may require customized arrangements with Xiaomi.
Use Chat Completions when you need the broadly documented messages-based interface. Use the Responses-compatible endpoint when your integration or tool ecosystem specifically expects the newer Responses pattern, including some coding and agent tools.
The relevant endpoints are /chat/completions, /responses, and /models. MiMo V2.5 and MiMo V2.5 Pro remain present in some documentation but are scheduled for deprecation on October 21, 2026, at 10:00 Beijing time. They are not appropriate starting points for new production work when V2.6 equivalents are available.
Make your first API request
The following Python example uses the current OpenAI Python SDK with Xiaomi's OpenAI-compatible base URL. Install the SDK with pip install -U openai, set MIMO_API_KEY, and then run the request.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MIMO_API_KEY"],
base_url="https://api.xiaomimimo.com/v1",
)
response = client.chat.completions.create(
model="mimo-v2.6-flash",
messages=[
{"role": "user", "content": "Explain API keys in one sentence."}
],
max_completion_tokens=256,
)
print(response.choices[0].message.content)The same request can be made directly over HTTP:
curl https://api.xiaomimimo.com/v1/chat/completions
-H "api-key: $MIMO_API_KEY"
-H "Content-Type: application/json"
-d '{
"model": "mimo-v2.6-flash",
"messages": [
{"role": "user", "content": "Explain API keys in one sentence."}
],
"max_completion_tokens": 256
}'Understanding the response
A Chat Completions response contains a choices array. The generated assistant text is normally available at choices[0].message.content. Applications should still handle cases where the model returns a tool call instead of text, or where the response contains an error, incomplete output, or no usable content.
For production code, check the HTTP status, handle malformed responses, and validate any content that will be used to control another system. Do not treat model-generated text or tool arguments as automatically trusted input.
How Xiaomi MiMo API pricing works
MiMo supports pay-as-you-go billing based on token usage. Input tokens, output tokens, and cached input tokens have separate rates, and cache hits cost less than ordinary input processing. Xiaomi documents domestic RMB and overseas dollar pricing separately. Web search is billed separately from model tokens.
As of September 22, 2026, documented overseas real-time rates are:
| Model | Cached input | Input | Output |
|---|---|---|---|
| MiMo V2.6 Pro | $0.0036 per million tokens | $0.435 per million tokens | $0.87 per million tokens |
| MiMo V2.6 Flash | $0.0028 per million tokens | $0.14 per million tokens | $0.28 per million tokens |
Cache-writing charges are documented as temporarily free. Batch inference costs 50% of the corresponding real-time rates for supported models. The platform also offers fixed Token Plan packages with separate quotas, credentials, and base URLs. A Token Plan balance cannot be assumed to work like ordinary pay-as-you-go account credit.
Actual application cost depends on prompt length, generated output, cache behavior, model selection, web-search use, and the amount of work sent through batch processing. Measure token usage with representative requests before setting a budget.
What can developers build with the API?
Streaming responses
Streaming uses server-sent events to deliver output incrementally instead of waiting for the complete response. This is useful for chat interfaces and long-running generations because the user can see partial text sooner. Streamed events may include generated text, reasoning-related fields where applicable, and tool-call deltas.
stream = client.chat.completions.create(
model="mimo-v2.6-flash",
messages=[
{"role": "user", "content": "List three API testing tips."}
],
stream=True,
max_completion_tokens=256,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Function and tool calling
Function calling lets you describe an application function with a JSON Schema-style parameter definition. The model can request that function, your application executes it, and your application sends the result back in a later conversation turn. The model does not execute your function by itself.
response = client.chat.completions.create(
model="mimo-v2.6-pro",
messages=[
{"role": "user", "content": "What is the weather in Seattle?"}
],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}],
tool_choice="auto",
max_completion_tokens=512,
)
message = response.choices[0].message
if message.tool_calls:
for call in message.tool_calls:
print(call.function.name, call.function.arguments)Validate the function name and arguments before executing anything. For tools that affect accounts, payments, files, or other external systems, add application-level authorization and confirmation checks.
Structured and JSON output
Chat Completions supports JSON mode with response_format: {"type": "json_object"}. You must also instruct the model to return JSON. JSON mode does not remove the need for application validation: parse the response and validate it against your own expected schema before storing or acting on it.
response = client.chat.completions.create(
model="mimo-v2.6-flash",
messages=[
{
"role": "system",
"content": "Return only valid JSON with the keys answer and confidence."
},
{"role": "user", "content": "Explain what an API key is."}
],
response_format={"type": "json_object"},
max_completion_tokens=256,
)
print(response.choices[0].message.content)Web search
The Web Search Plugin lets supported models retrieve current public web information. It must be enabled in the console and is charged separately. It can be used with streaming or non-streaming responses and can be combined with custom function tools.
Because search results can change and may be incomplete, applications should decide how to display sources, handle no-result cases, and distinguish retrieved information from the model's own generated explanation.
Image, audio, video, and file input
Supported MiMo models accept image input through publicly accessible image URLs or Base64 data URLs. Documented image formats include JPEG, PNG, GIF, WebP, and BMP, with a single URL or encoded image limited to 50 MB.
The documented real-time multimodal models do not support local image-file upload through the listed image-understanding interface. Convert an image to an accepted data URL or make it available through an accessible URL where appropriate. The platform also documents audio and video input for supported models, but availability is model-specific.
File upload is available for Batch Inference. Developers upload JSONL job files, create an asynchronous batch job, and later download result or error files. Batch input and output files are retained for 30 days by default and then automatically deleted.
SDKs, playground, and developer tools
There is no Xiaomi-specific SDK verified in the supplied documentation. The official examples use the OpenAI Python SDK and the Anthropic Python SDK through compatible base URLs. JavaScript and PHP applications can use the OpenAI-compatible HTTP API or a compatible client library.
The MiMo platform console provides API-key management, model access, billing controls, and activation for features such as web search. A platform playground is available at https://platform.xiaomimimo.com/. Xiaomi also provides a separate Agent ecosystem with MCP, Skills, and agent publication. That ecosystem is distinct from a persistent assistant-object API: the MiMo API provides model-level tool calling and Responses compatibility, but the reviewed documentation does not describe a persistent assistant resource.
Limits and production considerations
For MiMo V2.6 Pro and Flash, the documented limit is 100 requests per minute and 10 million tokens per minute per account and model. The limit aggregates requests across API keys belonging to the same account. ASR and TTS models are documented at 100 requests per minute. UltraSpeed uses customized service arrangements rather than a published standard quota.
Xiaomi does not publish a general latency SLA in the reviewed API documentation. High server load can cause delays or HTTP 429 responses. Production clients should use request timeouts, exponential backoff, bounded retry counts, and idempotent application-level handling so a retry does not accidentally repeat an external action.
Keep separate credentials and configuration for pay-as-you-go, Token Plans, and regional batch services. Confirm model-specific support before relying on multimodal input, web search, audio, video, or structured output. Also note that the public documentation reviewed here does not establish a general retention period for ordinary real-time prompts and outputs or a universal training-use policy. Batch files have the documented 30-day retention period, but developers should obtain current contractual or policy confirmation for other data-handling requirements.
When is Xiaomi MiMo API a good choice?
MiMo is a sensible choice when you want an OpenAI-compatible integration, current MiMo V2.6 models, token-based billing, image understanding, tool calling, web search, structured JSON, or batch inference in one Xiaomi platform. Compatibility with familiar SDKs can reduce the amount of new client code required.
It may be a poor fit if you require a globally uniform service with identical model features in every region, a published general latency SLA, a mature persistent-assistant resource model, or clearly documented real-time data-retention and training guarantees. The separate credential types, regional batch endpoint, model-specific capabilities, and scheduled V2.5 deprecation also require operational planning.
For a new project, begin with mimo-v2.6-flash for cost-conscious workloads or mimo-v2.6-pro when the higher-capability option is more appropriate. Test the exact prompts, tools, input formats, limits, and billing path your application will use before moving to production.
