What is the MiniMax API?
MiniMax Open Platform is a developer service for adding generative AI features to applications. Its APIs cover language models, speech synthesis, image generation, video generation, music-related services, file management, and server-side tools such as web search. It is separate from MiniMax consumer applications and is intended for developers who need programmatic access, usage-based billing, and application-controlled workflows.
For language applications, MiniMax currently recommends its Anthropic-compatible API. An OpenAI-compatible API is also available, which can reduce migration work for applications already built with the OpenAI SDK. Native HTTP, multipart-upload, asynchronous-task, and WebSocket interfaces are used for media and file workflows.
The platform is suitable for chat applications, coding assistants, agent workflows, document and video analysis, speech interfaces, content-generation tools, and applications that combine text with image or video input. It is less suitable when an application requires a single uniform product experience, guaranteed unrestricted free usage, or a clearly documented hosted assistants system with persistent assistant, thread, and run objects.
Getting access and creating an API key
Start by registering or signing in to the MiniMax API Platform. Create an API key in the developer console, then add account balance or another eligible resource before sending paid requests. Standard Open Platform API keys are separate from Token Plan subscription keys, so a subscription key should not be assumed to work with every pay-as-you-go endpoint.
Send the key as a bearer token in the Authorization header. Store it in an environment variable or a server-side secret manager. Do not place it in browser JavaScript, mobile application bundles, public repositories, or client-side HTML, because users could extract and misuse it.
export MINIMAX_API_KEY="your-api-key"
curl --fail-with-body --silent --show-error
https://api.minimax.io/v1/chat/completions
-H "Authorization: Bearer ${MINIMAX_API_KEY}"
-H "Content-Type: application/json"The international API base URL is https://api.minimax.io. Regional endpoint and account requirements should be checked before deployment because model availability, payment options, and access rules can vary by region.
Choosing an interface and model
For a new language integration, the Anthropic-compatible API is the recommended starting point. The OpenAI-compatible interface is a practical alternative for existing OpenAI SDK applications and uses the https://api.minimax.io/v1 base URL. The Anthropic-compatible endpoint uses https://api.minimax.io/anthropic.
MiniMax-M3 is the current flagship language model listed in the platform documentation. It is positioned for long-context work, reasoning, coding, agentic workflows, tool use, and multimodal input. The platform also lists MiniMax-M2.7, MiniMax-M2.7-highspeed, MiniMax-M2.5, MiniMax-M2.5-highspeed, MiniMax-M2.1, MiniMax-M2.1-highspeed, and MiniMax-M2.
Choose M3 when the application needs the current flagship language capabilities, long context, or multimodal text, image, and video input. Compare the M2.x and high-speed variants when throughput, price, or response speed matters more than using the newest model. Model names, limits, and availability can change, so production systems should verify the current documentation rather than hard-code assumptions about model support.
Making a first language-model request
The OpenAI-compatible API accepts ordinary chat messages. The following example uses the current OpenAI Python SDK generation with MiniMax's compatible base URL.
pip install openaiimport os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MINIMAX_API_KEY"],
base_url="https://api.minimax.io/v1",
)
response = client.chat.completions.create(
model="MiniMax-M3",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain API streaming in three sentences."},
],
extra_body={"thinking": {"type": "disabled"}},
max_completion_tokens=512,
)
print(response.choices[0].message.content)The request specifies a model and a sequence of messages. The response contains a list of choices; for a normal single-answer request, the generated text is available at response.choices[0].message.content. The thinking setting shown here follows the supplied MiniMax OpenAI-compatible example and disables extended thinking for a short response.
How MiniMax API pricing works
Standard API access uses pay-as-you-go billing. Language models are generally charged per million input and output tokens, with separate prompt-cache pricing where applicable. Other services use their own units, including characters or audio hours for speech, output seconds for video, images for image generation, and requests or tasks for some platform features.
The documented MiniMax-M3 standard price is $0.30 per million input tokens and $1.20 per million output tokens for requests with up to 512,000 input tokens. Requests above that input size are listed at $0.60 per million input tokens and $2.40 per million output tokens. Prompt-cache reads are priced separately.
A priority service tier is available for supported language requests at 1.5 times standard pricing. It is intended to improve admission priority, response speed, and request reliability; it does not change the underlying application design requirements.
Track usage by unit rather than treating every request as equivalent. A language request may consume input and output tokens, while a speech, image, or video workflow may incur character, image, audio-hour, output-second, request, or task charges.
Streaming, tools, and multimodal input
Streaming responses
Streaming sends generated output incrementally instead of waiting for the complete response. It is useful for chat interfaces because the application can display text as it arrives. The OpenAI-compatible API uses the standard stream parameter.
stream = client.chat.completions.create(
model="MiniMax-M3",
messages=[
{"role": "user", "content": "Explain function calling briefly."}
],
stream=True,
extra_body={"thinking": {"type": "disabled"}},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()Speech synthesis also supports HTTP and WebSocket streaming, allowing audio to be produced and consumed incrementally.
Function and tool calling
MiniMax-M3 supports function tools through the OpenAI-compatible interface. A tool definition describes a function that your application owns, such as a weather lookup or database query. The model can request the function, but your server must validate the arguments, execute the function, and send the result back.
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}]
response = client.chat.completions.create(
model="MiniMax-M3",
messages=[{"role": "user", "content": "What is the weather in Boston?"}],
tools=tools,
tool_choice="auto",
extra_body={"thinking": {"type": "disabled"}},
)
print(response.choices[0].message)In a multi-turn tool workflow, preserve the complete assistant message, including tool-call information and any required reasoning fields, before appending tool results. Omitting that information can break the model's tool-call continuity.
Image, video, and file input
MiniMax-M3 supports text, image, and video input through the OpenAI-compatible message format. Images can be supplied by URL or base64 data. Larger videos can be uploaded through the Files API and referenced with an mm_file:// identifier.
The File Management API supports uploading, listing, retrieving, reading, and deleting files. Depending on the endpoint, files can be used for video understanding, video-generation inputs, voice cloning, prompt audio, and asynchronous long-text speech synthesis. Temporary files have validity limits: video-understanding and video-generation input files are documented as valid for up to seven days, so applications should be prepared to upload them again.
Media generation and server-side tools
Media APIs do not always follow the same request-and-response pattern as chat. Video generation is asynchronous: create a task, retain its task identifier, query its status, and retrieve the output URL after completion. Long-text speech synthesis also provides an asynchronous workflow for requests beyond the synchronous character limit.
The platform provides a server-side web_search tool through the Anthropic Messages API and the OpenAI Responses API. This lets the model search current web information while producing a response, with separate per-request pricing. MiniMax also publishes MCP implementations for connecting media-generation capabilities to MCP-compatible agent environments.
The public API documentation does not establish a separate persistent Assistants API with hosted assistants, threads, and runs. If your application needs those objects, implement conversation storage, tool execution, and task orchestration in your own application unless a separately documented MiniMax product supplies them.
SDKs, HTTP access, and the playground
MiniMax documents the Anthropic SDK as the recommended language integration for language models and also provides current OpenAI SDK examples for Python and Node.js. The platform is HTTP-based, so PHP and other languages can call it directly with standard HTTP libraries. Native media endpoints may require JSON requests, multipart uploads, asynchronous polling, or WebSocket connections.
The Open Platform console provides a playground for trying the service and managing API access. Use it to validate credentials and explore requests, but reproduce important tests in code before production because quotas, model availability, billing, and regional access can differ.
Limits and production considerations
Published limits vary by model and endpoint. The documentation lists MiniMax-M3 at 200 requests per minute and 10,000,000 tokens per minute. The listed M2.x language models generally allow 500 requests per minute and 20,000,000 tokens per minute. Media services have separate connection, request, and concurrent-task limits; for example, the documented H3 video-generation limits include 300 requests per minute and 30 maximum in-flight tasks.
Published approximate output speeds include at least 100 tokens per second for M3, about 60 tokens per second for M2.7 and M2.5, and about 100 tokens per second for high-speed variants. These are estimates rather than latency guarantees. Prompt size, traffic, region, streaming, service tier, and workload affect actual performance.
- Use exponential backoff for transient errors and respect both requests-per-minute and tokens-per-minute limits.
- Track language tokens separately from speech characters, audio hours, image counts, video seconds, requests, and asynchronous tasks.
- Use streaming for interactive text or speech experiences.
- Keep uploaded-file identifiers and media download URLs temporary unless the relevant documentation says otherwise.
- Keep regional endpoint selection, billing account, and API key configuration aligned.
- Do not send confidential or regulated data without reviewing the current API agreement and privacy policy.
Privacy and data handling
The MiniMax API privacy policy states that personal data may be retained as long as necessary or permitted by law or to fulfill the purpose for which it was collected. It also states that personal data may be stored in data centers in the United States and processed by MiniMax or third-party vendors subject to applicable law.
The policy says that MiniMax does not use input personal data to infer characteristics about an individual or use personal data for training to profile or target consumers. It does not provide a simple blanket statement that every API prompt and output is excluded from every form of model training. Review the current policy and obtain appropriate contractual assurances before sending sensitive production data.
When MiniMax API is a good or poor choice
MiniMax is a good choice when an application needs one developer platform covering language, coding, reasoning, multimodal input, speech, image, video, files, and web search. The OpenAI-compatible interface can simplify adoption for existing SDK-based applications, while the Anthropic-compatible interface is the platform's recommended language integration. M3's long context and tool-use support are useful for document analysis and agent workflows.
It may be a poor choice when the project requires identical availability and pricing in every country, a stable single-product API across all modalities, mature hosted assistant objects, unrestricted free usage, or a blanket no-training guarantee for submitted data. The ecosystem is divided across endpoint types and product-specific policies, and models, quotas, regional access, and media availability can change. Validate the exact model, endpoint, retention terms, and price before committing to a production architecture.
