Developer platform

Xiaomi MiMo API Open Platform

Developer overview for Xiaomi HyperAI, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API OpenAI-compatible Chat Completions API and Responses API
SDK support Official examples use the OpenAI Python SDK and Anthropic Python SDK through compatible base URLs. JavaScript and PHP developers can use the OpenAI-compatible HTTP API; no Xiaomi-specific SDK was verified.
Rate limits MiMo V2.6 Pro and Flash: 100 RPM and 10 million TPM per account and model. ASR and TTS models: 100 RPM. UltraSpeed has customized service availability. Limits aggregate all API keys under an account; 429 responses may occur during high load.
Platform

API overview

Endpoints

API access

Base URL https://api.xiaomimimo.com/v1
Primary API OpenAI-compatible Chat Completions API and Responses API
Pricing

API pricing

Pricing model Pay-as-you-go token billing, fixed Token Plan subscriptions, and discounted Batch API billing

Overseas real-time pricing includes MiMo V2.6 Pro at $0.435 per million input tokens and $0.87 per million output tokens, and MiMo V2.6 Flash at $0.14 per million input tokens and $0.28 per million output tokens. Cached-input pricing is lower. Batch infer

Developer experience

SDKs & usability

SDKs Official examples use the OpenAI Python SDK and Anthropic Python SDK through compatible base URLs. JavaScript and PHP developers can use the OpenAI-compatible HTTP API; no Xiaomi-specific SDK was verified.
Ease of use Moderate to easy because the platform supports OpenAI and Anthropic-compatible protocols and works with common SDKs. Account, credential type, regional endpoint, model capability, and billing differences require attention.
Documentation Good and improving. Official documentation includes quickstarts, API references, model catalogs, examples, structured output, multimodal input, web search, batch processing, rate limits, error codes, and migration updates.
Latency No general public latency SLA was verified. Xiaomi documents possible delays during high server load and offers MiMo V2.6 Pro UltraSpeed as a customized low-latency option.
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

The current MiMo API supports Chat Completions, Responses API compatibility, streaming, multi-turn messages, developer and system instructions, function tools, web search, JSON mode, image URL and Base64 input, audio and video input on supported models, model listing, and asynchronous Batch API processing. MiMo V2.6 Pro and Flash are the recommended current models. MiMo V2.5 and V2.5 Pro remain accessible in some documentation but are scheduled for deprecation on October 21, 2026. Web search requires console activation and is billed separately. The platform provides a first-party Xiaomi Agent ecosystem with MCP, Skills, and Agent publication, but the MiMo API documentation does not describe a persistent assistant-object resource. Local image and audio file upload is not supported by the documented real-time multimodal models; image URLs and Base64 data are supported. Batch Inference supports file upload for JSONL jobs and asynchronous processing.

Examples

API examples

set -euo pipefail

: "${MIMO_API_KEY:?Set MIMO_API_KEY first}"

curl --fail-with-body --silent --show-error \
  https://api.xiaomimimo.com/v1/chat/completions \
  -H "api-key: ${MIMO_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mimo-v2.6-flash",
    "messages": [
      {
        "role": "system",
        "content": "You are a concise technical assistant. Return only valid JSON with the keys answer and confidence."
      },
      {
        "role": "user",
        "content": "Explain what an API key is in one sentence."
      }
    ],
    "response_format": {
      "type": "json_object"
    },
    "max_completion_tokens": 256
  }' | python3 -c 'import json,sys; print(json.load(sys.stdin)["choices"][0]["message"]["content"])'

# Streaming example:
curl --fail-with-body --silent --show-error \
  https://api.xiaomimimo.com/v1/chat/completions \
  -H "api-key: ${MIMO_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mimo-v2.6-pro",
    "messages": [{"role": "user", "content": "Write three short API design principles."}],
    "stream": true,
    "max_completion_tokens": 256
  }'
# Install: pip install -U openai
import json
import os
from openai import OpenAI

api_key = os.environ.get("MIMO_API_KEY")
if not api_key:
    raise RuntimeError("MIMO_API_KEY is not set")

client = OpenAI(
    api_key=api_key,
    base_url="https://api.xiaomimimo.com/v1",
)

try:
    response = client.chat.completions.create(
        model="mimo-v2.6-flash",
        messages=[
            {
                "role": "system",
                "content": "You are a concise technical assistant.",
            },
            {"role": "user", "content": "What is an API key?"},
        ],
        max_completion_tokens=256,
    )
    print(response.choices[0].message.content)

    tool_response = client.chat.completions.create(
        model="mimo-v2.6-pro",
        messages=[
            {"role": "system", "content": "Use the tool when current weather is requested."},
            {"role": "user", "content": "What is the weather in Seattle?"},
        ],
        tools=[
            {
                "type": "function",
                "function": {
                    "name": "get_weather",
                    "description": "Get current weather for a city.",
                    "parameters": {
                        "type": "object",
                        "properties": {"city": {"type": "string"}},
                        "required": ["city"],
                    },
                },
            }
        ],
        tool_choice="auto",
        max_completion_tokens=512,
    )
    message = tool_response.choices[0].message
    if message.tool_calls:
        for call in message.tool_calls:
            arguments = json.loads(call.function.arguments)
            print({"tool": call.function.name, "arguments": arguments})
    else:
        print(message.content)

    stream = client.chat.completions.create(
        model="mimo-v2.6-flash",
        messages=[
            {"role": "user", "content": "Give me three concise API testing tips."}
        ],
        stream=True,
        max_completion_tokens=256,
    )
    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            print(chunk.choices[0].delta.content, end="", flush=True)
    print()
except Exception as exc:
    print(f"MiMo API request failed: {exc}")
// Install: npm install openai
import OpenAI from "openai";

const apiKey = process.env.MIMO_API_KEY;
if (!apiKey) throw new Error("MIMO_API_KEY is not set");

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.xiaomimimo.com/v1",
});

try {
  const response = await client.chat.completions.create({
    model: "mimo-v2.6-flash",
    messages: [
      { role: "system", content: "Answer concisely and accurately." },
      { role: "user", content: "What is an API key?" },
    ],
    max_completion_tokens: 256,
  });
  console.log(response.choices[0]?.message?.content ?? "");

  const toolResponse = await client.chat.completions.create({
    model: "mimo-v2.6-pro",
    messages: [
      { role: "user", content: "Find the weather for Seattle using the available function." },
    ],
    tools: [
      {
        type: "function",
        function: {
          name: "get_weather",
          description: "Get current weather for a city.",
          parameters: {
            type: "object",
            properties: { city: { type: "string" } },
            required: ["city"],
          },
        },
      },
    ],
    tool_choice: "auto",
    max_completion_tokens: 512,
  });
  const toolCalls = toolResponse.choices[0]?.message?.tool_calls ?? [];
  for (const call of toolCalls) {
    console.log({ name: call.function.name, arguments: JSON.parse(call.function.arguments) });
  }

  const stream = await client.chat.completions.create({
    model: "mimo-v2.6-flash",
    messages: [{ role: "user", content: "List three API testing tips." }],
    stream: true,
    max_completion_tokens: 256,
  });
  for await (const chunk of stream) {
    process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
  }
  process.stdout.write("\n");
} catch (error) {
  console.error("MiMo API request failed:", error);
  process.exitCode = 1;
}
<?php
$apiKey = getenv('MIMO_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('MIMO_API_KEY is not set');
}

$url = 'https://api.xiaomimimo.com/v1/chat/completions';
$payload = [
    'model' => 'mimo-v2.6-flash',
    'messages' => [
        [
            'role' => 'system',
            'content' => 'Answer concisely and accurately.'
        ],
        [
            'role' => 'user',
            'content' => 'What is an API key?'
        ]
    ],
    'max_completion_tokens' => 256
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'api-key: ' . $apiKey,
        'Content-Type: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 60,
]);

$raw = curl_exec($ch);
if ($raw === false) {
    throw new RuntimeException('cURL error: ' . curl_error($ch));
}
$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($raw, true);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? $raw;
    throw new RuntimeException("MiMo API HTTP {$status}: {$message}");
}
if (!is_array($data)) {
    throw new RuntimeException('Invalid JSON response');
}

$result = $data['choices'][0]['message']['content'] ?? null;
if ($result === null) {
    throw new RuntimeException('No assistant content in response: ' . $raw);
}
echo $result . PHP_EOL;
Policies

Data & usage

Data training

No specific official MiMo API statement confirming that real-time API inputs are or are not used for model training was found in the reviewed current documentation. Developers should not assume a no-training guarantee without obtaining current contractual or policy confirmation from Xiaomi.

Data retention

For Batch Inference, input and result files are retained for 30 days by default and automatically deleted after expiry. A separate public retention period for ordinary real-time API prompts and outputs was not verified in the reviewed documentation.

Rate limits

MiMo V2.6 Pro and Flash: 100 RPM and 10 million TPM per account and model. ASR and TTS models: 100 RPM. UltraSpeed has customized service availability. Limits aggregate all API keys under an account; 429 responses may occur during high load.

Developer guide

Xiaomi MiMo API: Beginner's Guide to Models, Pricing, and Integration

Xiaomi MiMo API Open Platform is Xiaomi's developer platform for accessing MiMo models through OpenAI- and Anthropic-compatible APIs. Developers can use API keys, Chat Completions, Responses API compatibility, multimodal inputs, streaming, function calling, web search, structured JSON output, speech models, and asynchronous batch inference. The current recommended models are MiMo V2.6 Pro and MiMo V2.6 Flash.
Xiaomi MiMo API Open Platform gives developers programmatic access to Xiaomi's MiMo models without requiring a Xiaomi-specific SDK. Its OpenAI-compatible interface supports familiar Chat Completions and Responses API patterns, while the platform adds image understanding, streaming, tool calling, web search, structured output, and batch processing. This guide explains how to obtain access, make a first request, estimate costs, choose models, and account for production limitations.
Xiaomi MiMo API Open Platform provides OpenAI- and Anthropic-compatible access to MiMo V2.6 models. Developers can use token billing, streaming, structured JSON output, function tools, web search, image understanding, and asynchronous batch inference, with separate credentials and quotas for some services.

What is the Xiaomi MiMo API?

The Xiaomi MiMo API Open Platform is Xiaomi's current developer-facing model-inference platform. It lets applications send prompts and other supported inputs to MiMo models and receive generated responses over HTTPS.

The platform is designed to be familiar to developers who have used other model APIs. Its primary interface is OpenAI-compatible, and Xiaomi also documents Anthropic-compatible access. The main real-time base URL is https://api.xiaomimimo.com/v1.

Use the API when you want to add model-generated text, image understanding, tool use, web-assisted answers, structured JSON, or asynchronous batch processing to an application. It is aimed at developers building services, automation, agent workflows, coding tools, and other software integrations rather than people looking for a standalone consumer chatbot.

How to get access and create an API key

Start by creating or using a Xiaomi account and opening the MiMo API Open Platform console. From the console, create an API key and check which billing mode and regional service are associated with it.

Pay-as-you-go API keys use the sk- format. Token Plan credentials use separate tp- or team-plan formats and must be used with their dedicated base URL. These credential types and balances are not interchangeable.

Send the key with either the api-key header or a Bearer authorization header, depending on the integration pattern. Store it in an environment variable or secret manager rather than placing it in browser code, a source repository, or a mobile application.

Xiaomi's current documentation primarily describes personal-account login, and some services may require additional account or real-name verification. Check the console before planning an automated production signup process.

Which model and API should you choose?

For new integrations, start with the current MiMo V2.6 family:

  • mimo-v2.6-pro is the general higher-capability option and supports full-modality use cases.
  • mimo-v2.6-flash is the lower-cost option for many text and multimodal workloads.
  • mimo-v2.6-pro-ultraspeed is intended for latency-sensitive workloads, but its availability and service terms may require customized arrangements with Xiaomi.

Use Chat Completions when you need the broadly documented messages-based interface. Use the Responses-compatible endpoint when your integration or tool ecosystem specifically expects the newer Responses pattern, including some coding and agent tools.

The relevant endpoints are /chat/completions, /responses, and /models. MiMo V2.5 and MiMo V2.5 Pro remain present in some documentation but are scheduled for deprecation on October 21, 2026, at 10:00 Beijing time. They are not appropriate starting points for new production work when V2.6 equivalents are available.

Make your first API request

The following Python example uses the current OpenAI Python SDK with Xiaomi's OpenAI-compatible base URL. Install the SDK with pip install -U openai, set MIMO_API_KEY, and then run the request.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MIMO_API_KEY"],
    base_url="https://api.xiaomimimo.com/v1",
)

response = client.chat.completions.create(
    model="mimo-v2.6-flash",
    messages=[
        {"role": "user", "content": "Explain API keys in one sentence."}
    ],
    max_completion_tokens=256,
)

print(response.choices[0].message.content)

The same request can be made directly over HTTP:

curl https://api.xiaomimimo.com/v1/chat/completions 
  -H "api-key: $MIMO_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "mimo-v2.6-flash",
    "messages": [
      {"role": "user", "content": "Explain API keys in one sentence."}
    ],
    "max_completion_tokens": 256
  }'

Understanding the response

A Chat Completions response contains a choices array. The generated assistant text is normally available at choices[0].message.content. Applications should still handle cases where the model returns a tool call instead of text, or where the response contains an error, incomplete output, or no usable content.

For production code, check the HTTP status, handle malformed responses, and validate any content that will be used to control another system. Do not treat model-generated text or tool arguments as automatically trusted input.

How Xiaomi MiMo API pricing works

MiMo supports pay-as-you-go billing based on token usage. Input tokens, output tokens, and cached input tokens have separate rates, and cache hits cost less than ordinary input processing. Xiaomi documents domestic RMB and overseas dollar pricing separately. Web search is billed separately from model tokens.

As of September 22, 2026, documented overseas real-time rates are:

ModelCached inputInputOutput
MiMo V2.6 Pro$0.0036 per million tokens$0.435 per million tokens$0.87 per million tokens
MiMo V2.6 Flash$0.0028 per million tokens$0.14 per million tokens$0.28 per million tokens

Cache-writing charges are documented as temporarily free. Batch inference costs 50% of the corresponding real-time rates for supported models. The platform also offers fixed Token Plan packages with separate quotas, credentials, and base URLs. A Token Plan balance cannot be assumed to work like ordinary pay-as-you-go account credit.

Actual application cost depends on prompt length, generated output, cache behavior, model selection, web-search use, and the amount of work sent through batch processing. Measure token usage with representative requests before setting a budget.

What can developers build with the API?

Streaming responses

Streaming uses server-sent events to deliver output incrementally instead of waiting for the complete response. This is useful for chat interfaces and long-running generations because the user can see partial text sooner. Streamed events may include generated text, reasoning-related fields where applicable, and tool-call deltas.

stream = client.chat.completions.create(
    model="mimo-v2.6-flash",
    messages=[
        {"role": "user", "content": "List three API testing tips."}
    ],
    stream=True,
    max_completion_tokens=256,
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Function and tool calling

Function calling lets you describe an application function with a JSON Schema-style parameter definition. The model can request that function, your application executes it, and your application sends the result back in a later conversation turn. The model does not execute your function by itself.

response = client.chat.completions.create(
    model="mimo-v2.6-pro",
    messages=[
        {"role": "user", "content": "What is the weather in Seattle?"}
    ],
    tools=[{
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather for a city.",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"]
            }
        }
    }],
    tool_choice="auto",
    max_completion_tokens=512,
)

message = response.choices[0].message
if message.tool_calls:
    for call in message.tool_calls:
        print(call.function.name, call.function.arguments)

Validate the function name and arguments before executing anything. For tools that affect accounts, payments, files, or other external systems, add application-level authorization and confirmation checks.

Structured and JSON output

Chat Completions supports JSON mode with response_format: {"type": "json_object"}. You must also instruct the model to return JSON. JSON mode does not remove the need for application validation: parse the response and validate it against your own expected schema before storing or acting on it.

response = client.chat.completions.create(
    model="mimo-v2.6-flash",
    messages=[
        {
            "role": "system",
            "content": "Return only valid JSON with the keys answer and confidence."
        },
        {"role": "user", "content": "Explain what an API key is."}
    ],
    response_format={"type": "json_object"},
    max_completion_tokens=256,
)

print(response.choices[0].message.content)

The Web Search Plugin lets supported models retrieve current public web information. It must be enabled in the console and is charged separately. It can be used with streaming or non-streaming responses and can be combined with custom function tools.

Because search results can change and may be incomplete, applications should decide how to display sources, handle no-result cases, and distinguish retrieved information from the model's own generated explanation.

Image, audio, video, and file input

Supported MiMo models accept image input through publicly accessible image URLs or Base64 data URLs. Documented image formats include JPEG, PNG, GIF, WebP, and BMP, with a single URL or encoded image limited to 50 MB.

The documented real-time multimodal models do not support local image-file upload through the listed image-understanding interface. Convert an image to an accepted data URL or make it available through an accessible URL where appropriate. The platform also documents audio and video input for supported models, but availability is model-specific.

File upload is available for Batch Inference. Developers upload JSONL job files, create an asynchronous batch job, and later download result or error files. Batch input and output files are retained for 30 days by default and then automatically deleted.

SDKs, playground, and developer tools

There is no Xiaomi-specific SDK verified in the supplied documentation. The official examples use the OpenAI Python SDK and the Anthropic Python SDK through compatible base URLs. JavaScript and PHP applications can use the OpenAI-compatible HTTP API or a compatible client library.

The MiMo platform console provides API-key management, model access, billing controls, and activation for features such as web search. A platform playground is available at https://platform.xiaomimimo.com/. Xiaomi also provides a separate Agent ecosystem with MCP, Skills, and agent publication. That ecosystem is distinct from a persistent assistant-object API: the MiMo API provides model-level tool calling and Responses compatibility, but the reviewed documentation does not describe a persistent assistant resource.

Limits and production considerations

For MiMo V2.6 Pro and Flash, the documented limit is 100 requests per minute and 10 million tokens per minute per account and model. The limit aggregates requests across API keys belonging to the same account. ASR and TTS models are documented at 100 requests per minute. UltraSpeed uses customized service arrangements rather than a published standard quota.

Xiaomi does not publish a general latency SLA in the reviewed API documentation. High server load can cause delays or HTTP 429 responses. Production clients should use request timeouts, exponential backoff, bounded retry counts, and idempotent application-level handling so a retry does not accidentally repeat an external action.

Keep separate credentials and configuration for pay-as-you-go, Token Plans, and regional batch services. Confirm model-specific support before relying on multimodal input, web search, audio, video, or structured output. Also note that the public documentation reviewed here does not establish a general retention period for ordinary real-time prompts and outputs or a universal training-use policy. Batch files have the documented 30-day retention period, but developers should obtain current contractual or policy confirmation for other data-handling requirements.

When is Xiaomi MiMo API a good choice?

MiMo is a sensible choice when you want an OpenAI-compatible integration, current MiMo V2.6 models, token-based billing, image understanding, tool calling, web search, structured JSON, or batch inference in one Xiaomi platform. Compatibility with familiar SDKs can reduce the amount of new client code required.

It may be a poor fit if you require a globally uniform service with identical model features in every region, a published general latency SLA, a mature persistent-assistant resource model, or clearly documented real-time data-retention and training guarantees. The separate credential types, regional batch endpoint, model-specific capabilities, and scheduled V2.5 deprecation also require operational planning.

For a new project, begin with mimo-v2.6-flash for cost-conscious workloads or mimo-v2.6-pro when the higher-capability option is more appropriate. Test the exact prompts, tools, input formats, limits, and billing path your application will use before moving to production.

Sources 13