Developer platform

MiniMax Open Platform API

Developer overview for MiniMax, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API MiniMax API Platform with Anthropic-compatible and OpenAI-compatible language APIs
SDK support Official documentation supports the Anthropic SDK as the recommended language integration and the OpenAI SDK for Python and Node.js. Raw HTTP is available for PHP and other languages. Official MiniMax MCP implementations are available in Python and JavaSc
Rate limits Published limits vary by API and model. MiniMax-M3: 200 RPM and 10,000,000 TPM. MiniMax-M2.7, M2.7-highspeed, M2.5, M2.5-highspeed, M2.1, M2.1-highspeed, and M2: generally 500 RPM and 20,000,000 TPM. H3 video generation: 300 RPM and 30 maximum inflight ta
Platform

API overview

Endpoints

API access

Base URL https://api.minimax.io
Primary API MiniMax API Platform with Anthropic-compatible and OpenAI-compatible language APIs
Pricing

API pricing

Pricing model Usage-based pay-as-you-go with token, character, image, audio, video-second, request, and task-based pricing; separate Token Plan and credit options are also available.

MiniMax-M3 standard pricing is listed at $0.30 per million input tokens and $1.20 per million output tokens for requests up to 512k input tokens; larger-context and priority pricing is higher. Other modalities use separate unit prices.

Developer experience

SDKs & usability

SDKs Official documentation supports the Anthropic SDK as the recommended language integration and the OpenAI SDK for Python and Node.js. Raw HTTP is available for PHP and other languages. Official MiniMax MCP implementations are available in Python and JavaSc
Ease of use Good. Standard HTTP APIs are available, and language models can be accessed with the official OpenAI SDK or the recommended Anthropic SDK. Media APIs use native HTTP, multipart, asynchronous-task, or WebSocket patterns.
Documentation Good and broad, with current API-overview, model, pricing, rate-limit, SDK, file, media, and server-tool documentation. Some capabilities are documented across separate endpoint pages and availability can vary by model or region.
Latency Published approximate output speeds include 100+ tokens per second for MiniMax-M3, about 60 tokens per second for M2.7 and M2.5, and about 100 tokens per second for high-speed variants. Actual latency varies by workload, context size, traffic, region, str
Features

API capabilities

✓ Streaming
✓ Function calling
✓ File uploads
✓ Web search
✓ Image input
✓ Playground
Feature notes

The platform provides language, speech, video, image, music, file-management, and server-tool APIs. MiniMax-M3 is the current flagship language model listed in the API documentation and supports a one-million-token total context window, text/image/video input, streaming, tool calling, and reasoning controls. The recommended language integration is the Anthropic-compatible API, with OpenAI-compatible Chat Completions also supported. Server-side web_search is available through the Anthropic Messages API and OpenAI Responses API. Files can be uploaded for video understanding, video generation, voice cloning, prompt audio, and asynchronous long-text speech synthesis. Speech supports HTTP and WebSocket streaming. Video generation is asynchronous and uses task creation and status-query workflows. Music APIs have restrictions for new users as of August 20, 2026. The public documentation does not establish a separate persistent hosted Assistants API with assistant, thread, and run objects. Fine-tuning is not clearly documented as a current public API capability.

Examples

API examples

curl --fail-with-body --silent --show-error https://api.minimax.io/v1/chat/completions \
  -H "Authorization: Bearer ${MINIMAX_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMax-M3",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain how streaming responses work."}
    ],
    "thinking": {"type": "disabled"},
    "max_completion_tokens": 512
  }' | jq -r '.choices[0].message.content'

curl --fail-with-body --silent --show-error https://api.minimax.io/v1/chat/completions \
  -H "Authorization: Bearer ${MINIMAX_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMax-M3",
    "messages": [
      {"role": "user", "content": "What is the weather in Boston?"}
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_weather",
          "description": "Get current weather for a city",
          "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"]
          }
        }
      }
    ],
    "tool_choice": "auto"
  }' | jq '.choices[0].message'
# Install: pip install openai
import json
import os
from openai import OpenAI

api_key = os.environ["MINIMAX_API_KEY"]
client = OpenAI(
    api_key=api_key,
    base_url="https://api.minimax.io/v1",
)

try:
    response = client.chat.completions.create(
        model="MiniMax-M3",
        messages=[
            {"role": "system", "content": "You are a helpful technical assistant."},
            {"role": "user", "content": "Give me three practical uses for a multimodal API."},
        ],
        extra_body={"thinking": {"type": "disabled"}},
        max_completion_tokens=512,
    )
    print(response.choices[0].message.content)
except Exception as exc:
    print(f"MiniMax API request failed: {exc}")

stream = client.chat.completions.create(
    model="MiniMax-M3",
    messages=[
        {"role": "system", "content": "Answer clearly and briefly."},
        {"role": "user", "content": "Explain function calling in one paragraph."},
    ],
    stream=True,
    extra_body={"thinking": {"type": "disabled"}},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()

def get_weather(city: str) -> dict:
    return {"city": city, "temperature_c": 18, "condition": "partly cloudy"}

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"]
        }
    }
}]

messages = [{"role": "user", "content": "What is the weather in Boston?"}]
first = client.chat.completions.create(
    model="MiniMax-M3",
    messages=messages,
    tools=tools,
    tool_choice="auto",
    extra_body={"thinking": {"type": "disabled"}},
)
assistant_message = first.choices[0].message
messages.append(assistant_message.model_dump(exclude_none=True))
if assistant_message.tool_calls:
    for call in assistant_message.tool_calls:
        arguments = json.loads(call.function.arguments)
        result = get_weather(arguments["city"])
        messages.append({
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result),
        })
    final = client.chat.completions.create(
        model="MiniMax-M3",
        messages=messages,
        extra_body={"thinking": {"type": "disabled"}},
    )
    print(final.choices[0].message.content)
else:
    print(assistant_message.content)
// Install: npm install openai
import OpenAI from "openai";

const apiKey = process.env.MINIMAX_API_KEY;
if (!apiKey) throw new Error("MINIMAX_API_KEY is required");

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.minimax.io/v1"
});

try {
  const response = await client.chat.completions.create({
    model: "MiniMax-M3",
    messages: [
      { role: "system", content: "You are a helpful assistant." },
      { role: "user", content: "Summarize the benefits of streaming APIs." }
    ],
    max_completion_tokens: 512,
    extra_body: { thinking: { type: "disabled" } }
  });
  console.log(response.choices[0].message.content);

  const stream = await client.chat.completions.create({
    model: "MiniMax-M3",
    messages: [
      { role: "user", content: "Explain multimodal input in two sentences." }
    ],
    stream: true,
    extra_body: { thinking: { type: "disabled" } }
  });
  for await (const chunk of stream) {
    const text = chunk.choices?.[0]?.delta?.content;
    if (text) process.stdout.write(text);
  }
  process.stdout.write("\n");
} catch (error) {
  console.error("MiniMax API request failed:", error instanceof Error ? error.message : error);
  process.exitCode = 1;
}
<?php
$apiKey = getenv('MINIMAX_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('MINIMAX_API_KEY is required');
}

$url = 'https://api.minimax.io/v1/chat/completions';
$payload = [
    'model' => 'MiniMax-M3',
    'messages' => [
        ['role' => 'system', 'content' => 'You are a helpful technical assistant.'],
        ['role' => 'user', 'content' => 'Explain how an API key should be stored securely.']
    ],
    'max_completion_tokens' => 512,
    'thinking' => ['type' => 'disabled']
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 60,
]);

$body = curl_exec($ch);
if ($body === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('cURL error: ' . $error);
}
$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    $message = $data['base_resp']['status_msg'] ?? $body;
    throw new RuntimeException('MiniMax API error (' . $status . '): ' . $message);
}

$result = $data['choices'][0]['message']['content'] ?? null;
if ($result === null) {
    throw new RuntimeException('The response did not contain message content.');
}

echo $result . PHP_EOL;
?>
Policies

Data & usage

Data training

The current MiniMax API Privacy Policy states that MiniMax does not use input personal data to infer characteristics about an individual or use personal data for training to profile or target consumers. It does not provide a simple blanket statement that all API prompts and outputs are excluded from all model-training processes. Developers should review the current policy and obtain appropriate contractual assurances for confidential, regulated, or sensitive workloads.

Data retention

Retention varies by data type and workflow. The privacy policy states that personal data may be retained for as long as necessary or permitted by law or to fulfill the purpose for which it was collected, after which it will be deleted or anonymized where applicable. Speech interfaces are described as stateless and do not store user data. Uploaded video-understanding files are retained for up to 7 days, and video-generation input files are valid for 7 days. Temporary voice-cloning and voice-design voices are deleted if unused for 168 hours.

Rate limits

Published limits vary by API and model. MiniMax-M3: 200 RPM and 10,000,000 TPM. MiniMax-M2.7, M2.7-highspeed, M2.5, M2.5-highspeed, M2.1, M2.1-highspeed, and M2: generally 500 RPM and 20,000,000 TPM. H3 video generation: 300 RPM and 30 maximum inflight ta

Developer guide

MiniMax API: Beginner's Guide to Models, Pricing, and Integration

MiniMax Open Platform is a pay-as-you-go developer API for language, speech, image, video, music, file-management, and web-search workflows. Developers can use the recommended Anthropic-compatible interface or the OpenAI-compatible API, authenticate with bearer API keys, stream responses, call tools, upload files, and build asynchronous media-generation workflows. MiniMax-M3 is the current flagship language model listed by the platform, while native endpoints handle capabilities such as speech synthesis and video generation.
MiniMax Open Platform gives developers API access to text generation, reasoning, coding, multimodal input, speech, image, video, file management, and server-side web search. A new integration normally starts with an Open Platform account, an API key, and either the recommended Anthropic-compatible language interface or the OpenAI-compatible Chat Completions API. This guide explains the practical setup, pricing model, capabilities, examples, limits, and trade-offs involved in using MiniMax in an application.
MiniMax Open Platform provides pay-as-you-go APIs for language, coding, reasoning, speech, image, video, files, and web search. Developers can use Anthropic-compatible or OpenAI-compatible language interfaces, authenticate with bearer keys, stream responses, call tools, upload files, and manage asynchronous media tasks. The guide covers setup, M3 and other models, pricing units, SDKs, limits, privacy considerations, and production trade-offs.

What is the MiniMax API?

MiniMax Open Platform is a developer service for adding generative AI features to applications. Its APIs cover language models, speech synthesis, image generation, video generation, music-related services, file management, and server-side tools such as web search. It is separate from MiniMax consumer applications and is intended for developers who need programmatic access, usage-based billing, and application-controlled workflows.

For language applications, MiniMax currently recommends its Anthropic-compatible API. An OpenAI-compatible API is also available, which can reduce migration work for applications already built with the OpenAI SDK. Native HTTP, multipart-upload, asynchronous-task, and WebSocket interfaces are used for media and file workflows.

The platform is suitable for chat applications, coding assistants, agent workflows, document and video analysis, speech interfaces, content-generation tools, and applications that combine text with image or video input. It is less suitable when an application requires a single uniform product experience, guaranteed unrestricted free usage, or a clearly documented hosted assistants system with persistent assistant, thread, and run objects.

Getting access and creating an API key

Start by registering or signing in to the MiniMax API Platform. Create an API key in the developer console, then add account balance or another eligible resource before sending paid requests. Standard Open Platform API keys are separate from Token Plan subscription keys, so a subscription key should not be assumed to work with every pay-as-you-go endpoint.

Send the key as a bearer token in the Authorization header. Store it in an environment variable or a server-side secret manager. Do not place it in browser JavaScript, mobile application bundles, public repositories, or client-side HTML, because users could extract and misuse it.

export MINIMAX_API_KEY="your-api-key"

curl --fail-with-body --silent --show-error 
  https://api.minimax.io/v1/chat/completions 
  -H "Authorization: Bearer ${MINIMAX_API_KEY}" 
  -H "Content-Type: application/json"

The international API base URL is https://api.minimax.io. Regional endpoint and account requirements should be checked before deployment because model availability, payment options, and access rules can vary by region.

Choosing an interface and model

For a new language integration, the Anthropic-compatible API is the recommended starting point. The OpenAI-compatible interface is a practical alternative for existing OpenAI SDK applications and uses the https://api.minimax.io/v1 base URL. The Anthropic-compatible endpoint uses https://api.minimax.io/anthropic.

MiniMax-M3 is the current flagship language model listed in the platform documentation. It is positioned for long-context work, reasoning, coding, agentic workflows, tool use, and multimodal input. The platform also lists MiniMax-M2.7, MiniMax-M2.7-highspeed, MiniMax-M2.5, MiniMax-M2.5-highspeed, MiniMax-M2.1, MiniMax-M2.1-highspeed, and MiniMax-M2.

Choose M3 when the application needs the current flagship language capabilities, long context, or multimodal text, image, and video input. Compare the M2.x and high-speed variants when throughput, price, or response speed matters more than using the newest model. Model names, limits, and availability can change, so production systems should verify the current documentation rather than hard-code assumptions about model support.

Making a first language-model request

The OpenAI-compatible API accepts ordinary chat messages. The following example uses the current OpenAI Python SDK generation with MiniMax's compatible base URL.

pip install openai
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MINIMAX_API_KEY"],
    base_url="https://api.minimax.io/v1",
)

response = client.chat.completions.create(
    model="MiniMax-M3",
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "Explain API streaming in three sentences."},
    ],
    extra_body={"thinking": {"type": "disabled"}},
    max_completion_tokens=512,
)

print(response.choices[0].message.content)

The request specifies a model and a sequence of messages. The response contains a list of choices; for a normal single-answer request, the generated text is available at response.choices[0].message.content. The thinking setting shown here follows the supplied MiniMax OpenAI-compatible example and disables extended thinking for a short response.

How MiniMax API pricing works

Standard API access uses pay-as-you-go billing. Language models are generally charged per million input and output tokens, with separate prompt-cache pricing where applicable. Other services use their own units, including characters or audio hours for speech, output seconds for video, images for image generation, and requests or tasks for some platform features.

The documented MiniMax-M3 standard price is $0.30 per million input tokens and $1.20 per million output tokens for requests with up to 512,000 input tokens. Requests above that input size are listed at $0.60 per million input tokens and $2.40 per million output tokens. Prompt-cache reads are priced separately.

A priority service tier is available for supported language requests at 1.5 times standard pricing. It is intended to improve admission priority, response speed, and request reliability; it does not change the underlying application design requirements.

Track usage by unit rather than treating every request as equivalent. A language request may consume input and output tokens, while a speech, image, or video workflow may incur character, image, audio-hour, output-second, request, or task charges.

Streaming, tools, and multimodal input

Streaming responses

Streaming sends generated output incrementally instead of waiting for the complete response. It is useful for chat interfaces because the application can display text as it arrives. The OpenAI-compatible API uses the standard stream parameter.

stream = client.chat.completions.create(
    model="MiniMax-M3",
    messages=[
        {"role": "user", "content": "Explain function calling briefly."}
    ],
    stream=True,
    extra_body={"thinking": {"type": "disabled"}},
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()

Speech synthesis also supports HTTP and WebSocket streaming, allowing audio to be produced and consumed incrementally.

Function and tool calling

MiniMax-M3 supports function tools through the OpenAI-compatible interface. A tool definition describes a function that your application owns, such as a weather lookup or database query. The model can request the function, but your server must validate the arguments, execute the function, and send the result back.

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"]
        }
    }
}]

response = client.chat.completions.create(
    model="MiniMax-M3",
    messages=[{"role": "user", "content": "What is the weather in Boston?"}],
    tools=tools,
    tool_choice="auto",
    extra_body={"thinking": {"type": "disabled"}},
)

print(response.choices[0].message)

In a multi-turn tool workflow, preserve the complete assistant message, including tool-call information and any required reasoning fields, before appending tool results. Omitting that information can break the model's tool-call continuity.

Image, video, and file input

MiniMax-M3 supports text, image, and video input through the OpenAI-compatible message format. Images can be supplied by URL or base64 data. Larger videos can be uploaded through the Files API and referenced with an mm_file:// identifier.

The File Management API supports uploading, listing, retrieving, reading, and deleting files. Depending on the endpoint, files can be used for video understanding, video-generation inputs, voice cloning, prompt audio, and asynchronous long-text speech synthesis. Temporary files have validity limits: video-understanding and video-generation input files are documented as valid for up to seven days, so applications should be prepared to upload them again.

Media generation and server-side tools

Media APIs do not always follow the same request-and-response pattern as chat. Video generation is asynchronous: create a task, retain its task identifier, query its status, and retrieve the output URL after completion. Long-text speech synthesis also provides an asynchronous workflow for requests beyond the synchronous character limit.

The platform provides a server-side web_search tool through the Anthropic Messages API and the OpenAI Responses API. This lets the model search current web information while producing a response, with separate per-request pricing. MiniMax also publishes MCP implementations for connecting media-generation capabilities to MCP-compatible agent environments.

The public API documentation does not establish a separate persistent Assistants API with hosted assistants, threads, and runs. If your application needs those objects, implement conversation storage, tool execution, and task orchestration in your own application unless a separately documented MiniMax product supplies them.

SDKs, HTTP access, and the playground

MiniMax documents the Anthropic SDK as the recommended language integration for language models and also provides current OpenAI SDK examples for Python and Node.js. The platform is HTTP-based, so PHP and other languages can call it directly with standard HTTP libraries. Native media endpoints may require JSON requests, multipart uploads, asynchronous polling, or WebSocket connections.

The Open Platform console provides a playground for trying the service and managing API access. Use it to validate credentials and explore requests, but reproduce important tests in code before production because quotas, model availability, billing, and regional access can differ.

Limits and production considerations

Published limits vary by model and endpoint. The documentation lists MiniMax-M3 at 200 requests per minute and 10,000,000 tokens per minute. The listed M2.x language models generally allow 500 requests per minute and 20,000,000 tokens per minute. Media services have separate connection, request, and concurrent-task limits; for example, the documented H3 video-generation limits include 300 requests per minute and 30 maximum in-flight tasks.

Published approximate output speeds include at least 100 tokens per second for M3, about 60 tokens per second for M2.7 and M2.5, and about 100 tokens per second for high-speed variants. These are estimates rather than latency guarantees. Prompt size, traffic, region, streaming, service tier, and workload affect actual performance.

  • Use exponential backoff for transient errors and respect both requests-per-minute and tokens-per-minute limits.
  • Track language tokens separately from speech characters, audio hours, image counts, video seconds, requests, and asynchronous tasks.
  • Use streaming for interactive text or speech experiences.
  • Keep uploaded-file identifiers and media download URLs temporary unless the relevant documentation says otherwise.
  • Keep regional endpoint selection, billing account, and API key configuration aligned.
  • Do not send confidential or regulated data without reviewing the current API agreement and privacy policy.

Privacy and data handling

The MiniMax API privacy policy states that personal data may be retained as long as necessary or permitted by law or to fulfill the purpose for which it was collected. It also states that personal data may be stored in data centers in the United States and processed by MiniMax or third-party vendors subject to applicable law.

The policy says that MiniMax does not use input personal data to infer characteristics about an individual or use personal data for training to profile or target consumers. It does not provide a simple blanket statement that every API prompt and output is excluded from every form of model training. Review the current policy and obtain appropriate contractual assurances before sending sensitive production data.

When MiniMax API is a good or poor choice

MiniMax is a good choice when an application needs one developer platform covering language, coding, reasoning, multimodal input, speech, image, video, files, and web search. The OpenAI-compatible interface can simplify adoption for existing SDK-based applications, while the Anthropic-compatible interface is the platform's recommended language integration. M3's long context and tool-use support are useful for document analysis and agent workflows.

It may be a poor choice when the project requires identical availability and pricing in every country, a stable single-product API across all modalities, mature hosted assistant objects, unrestricted free usage, or a blanket no-training guarantee for submitted data. The ecosystem is divided across endpoint types and product-specific policies, and models, quotas, regional access, and media availability can change. Validate the exact model, endpoint, retention terms, and price before committing to a production architecture.

Sources 11