MiniMax M2

MiniMax M2

by MiniMax · Legacy/open-weight model; current hosted availability and pricing are not confirmed in MiniMax's latest public model catalog

MiniMax M2 is an open-weight mixture-of-experts model focused on coding, reasoning, tool calling, and long-running agent workflows. It supports a 196,608-token context window and approximately 10 billion active parameters, while local deployment remains demanding because of its 230-billion total parameter count. Launch API pricing was $0.30 per million input tokens and $1.20 per million output tokens, but current hosted availability and pricing are unverified.

Text Reasoning Coding
MiniMax M2 is an open-weight language model from MiniMax built primarily for software development and agentic tasks. Its 196,608-token context window, interleaved reasoning, tool-calling support, and relatively small active parameter count make it suited to multi-step coding and research workflows, although its large total size makes self-hosting technically demanding.
Outputs

What MiniMax M2 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiniMax M2
Model type Coding
Context window 197K tokens
Release date 2025-10-27
Status Legacy/open-weight model; current hosted availability and pricing are not confirmed in MiniMax's latest public model catalog
Knowledge cutoff notes

No authoritative first-party knowledge-cutoff date was found for the exact MiniMax M2 model.

Model notes

MiniMax M2 is a 230B-total-parameter mixture-of-experts model with approximately 10B active parameters. It was released as an open-weight model for coding and agentic workflows and is available from the official MiniMaxAI/MiniMax-M2 repository on Hugging Face. Official materials describe interleaved thinking and tool workflows involving shell, browser, Python, retrieval, and MCP tools. The model has a 196,608-token maximum sequence length. The launch announcement listed API pricing of $0.30 per million input tokens and $1.20 per million output tokens, but current MiniMax catalog pages emphasize newer models and do not clearly confirm the original M2's present hosted availability. Local deployment requires substantial multi-GPU infrastructure despite the lower active-parameter count. The model is distributed under a modified MIT license.

Cost

Model pricing

Input $0.30 per 1 million input tokens at launch; current pricing unverified
Output $1.20 per 1 million output tokens at launch; current pricing unverified
Model guide

MiniMax M2: Open-Weight Coding Model for Long-Running Agents

MiniMax M2 is a 230-billion-parameter mixture-of-experts language model with approximately 10 billion active parameters, designed for coding, reasoning, tool use, and long-running agent workflows.

What is MiniMax M2?

MiniMax M2 is an open-weight mixture-of-experts language model released by MiniMax on October 27, 2025. It is designed chiefly for coding, reasoning, tool use, and agent workflows rather than for image, audio, or video generation.

The model has approximately 230 billion total parameters, but activates about 10 billion parameters for each token. This mixture-of-experts design can reduce the computation needed for each individual token compared with a dense model of the same total size. It does not, however, make the model small: local deployment still requires substantial GPU memory and generally multi-GPU infrastructure.

MiniMax M2 is distributed as an open-weight model through the official MiniMax-AI repository and Hugging Face model repository. The model is available under a modified MIT license, according to the supplied official materials.

Primary purpose and position in the MiniMax lineup

M2 occupies a specialist position in MiniMax's model portfolio. Its design emphasizes software engineering, tool-using agents, and long-running tasks that involve repeated planning and execution. It is not presented here as MiniMax's general consumer product or as a multimodal content-generation system.

MiniMax's current public catalog emphasizes newer models such as M3 and M2.7, so M2 should be treated as a legacy or independently useful open-weight model rather than assumed to be the provider's current flagship hosted option. Its open weights remain relevant for developers who want to inspect, adapt, or self-host the model, but its current hosted availability and pricing should be verified before building a production integration.

Core capabilities

M2 generates text and supports reasoning-oriented workflows. Its most relevant capabilities include:

  • Coding: software development, code repair, multi-file changes, and codebase analysis.
  • Reasoning: multi-step planning and interleaved thinking for tasks that require several intermediate decisions.
  • Tool use: workflows involving shell commands, browser interactions, Python execution, retrieval, and MCP tools.
  • Long-context processing: analysis of large codebases, documentation sets, and extended task histories.
  • Agent operation: repeated cycles of planning, tool calls, observation, and further action.

Interleaved thinking means that reasoning is integrated between ordinary responses and actions rather than treated as a single isolated step. MiniMax's materials indicate that thinking content represented with think tags should be retained in the conversational history when M2 is used locally or through compatible APIs. Implementations therefore need to follow the model's documented prompting and history conventions instead of silently discarding those sections.

Context window and output limits

The verified maximum sequence length for MiniMax M2 is 196,608 tokens. A token is a piece of text used by the model, so this limit includes the relevant input and conversation history as well as generated content, depending on the serving implementation.

This unusually large context is useful for tasks such as reviewing a substantial repository, keeping a long agent trace available, or combining source files with technical documentation. It does not guarantee that every detail in a very large prompt will receive equal attention, and it does not remove the need for sensible retrieval and context management.

A separate maximum output-token limit was not clearly documented in the supplied first-party materials. Developers should therefore avoid assuming that the full context window can be emitted as one response and should confirm the limit exposed by the selected serving framework or hosted endpoint.

Coding and agent workflows

M2 is most distinctive when it is connected to tools. The documented examples include terminal or shell operations, browser actions, Python execution, retrieval, and MCP-based tools. These capabilities allow an application to delegate more than code completion: the model can help inspect files, plan changes, run commands, examine results, and continue working through a multi-step task.

In practice, a coding agent built around M2 might use the model to:

  1. Inspect a repository and identify the files relevant to a bug.
  2. Propose a repair plan and request or execute tool actions.
  3. Modify several files while preserving the broader project structure.
  4. Run tests or scripts through a shell or Python tool.
  5. Interpret failures and revise the implementation.

Tool access does not mean the model independently guarantees safe execution. Shell commands, browser access, file changes, and code execution should be isolated and permissioned by the application. The model's ability to call tools is a capability of the model-and-harness combination, not proof that every deployment includes the same tools or safety controls.

Supported modalities and output types

MiniMax M2 is text-only in the supplied specification. It accepts text input and produces text output. The research does not verify native image, audio, or video input, and it does not describe native image, audio, video, music, speech, or embedding output for this model.

This makes M2 a poor fit when the central task is visual understanding, media generation, voice interaction, or multimodal content creation. It can still participate in workflows that call external tools for those tasks, but that is different from having those modalities built into the model itself.

Speed, cost, and deployment trade-offs

The approximately 10-billion active-parameter figure helps explain why M2 can be attractive for inference efficiency relative to a dense 230-billion-parameter model. The total model size remains large, however, so the active-parameter count should not be interpreted as a low hardware requirement.

At launch, MiniMax listed API pricing of $0.30 per million input tokens and $1.20 per million output tokens. These were launch prices, and the original hosted M2 pricing is not confirmed as current. MiniMax's newer catalog pages emphasize other models, so prospective users should verify whether M2 is still hosted, under what endpoint, and at what rate before estimating operating costs.

For local deployment, the trade-off is different. Open weights can provide control over hosting and experimentation, but the model's total size makes infrastructure more demanding than the active-parameter figure alone suggests. Official deployment guidance references frameworks including vLLM, SGLang, and MLX. Hardware requirements will depend on quantization, parallelism, context size, and serving configuration; the supplied research does not establish one universal memory requirement.

Strengths and limitations

Where M2 is strong

  • Specialization for coding agents: its intended workloads include code generation, code repair, tool use, and multi-step software tasks.
  • Large context: a 196,608-token maximum sequence length supports extensive code and task histories.
  • Open-weight access: developers can experiment with local deployment rather than relying exclusively on a hosted endpoint.
  • Tool-oriented reasoning: the model is designed for workflows that combine planning with shell, browser, Python, retrieval, and MCP actions.
  • Potential token-efficiency: the MoE architecture activates approximately 10 billion parameters per token rather than all 230 billion.

Where M2 is limited

  • Large deployment footprint: 230 billion total parameters make self-hosting demanding despite the lower active count.
  • Text-only operation: it is not a native image, audio, video, or music model.
  • Unclear current hosted status: present API availability and pricing are not confirmed by the supplied current catalog information.
  • Incomplete documented limits: the maximum output-token limit, knowledge cutoff, structured-output support, caching, batch API support, and fine-tuning status are not clearly established.
  • Operational safety requirements: applications must control permissions and validate tool actions rather than treating model-generated commands as automatically safe.

Best use cases

MiniMax M2 is a strong candidate for developers evaluating open-weight models for:

  • Multi-file software development and code repair.
  • Long-context codebase or documentation analysis.
  • Terminal, browser, retrieval, and code-execution agents.
  • Research assistants that need repeated tool calls and intermediate planning.
  • Self-hosted experiments with a large mixture-of-experts coding model.
  • Applications where open weights and deployment control matter more than a simple managed API.

It is less appropriate for a lightweight local assistant, a media-generation application, or a production system that requires a clearly documented current endpoint, output limit, structured-output guarantee, or lifecycle commitment.

When to choose MiniMax M2

Choose M2 when the central problem is text-based coding or agent automation and you can accommodate the infrastructure needed for a large open-weight model. It is especially appealing when a long context, tool-driven workflow, and the ability to experiment with model weights are more important than a turnkey hosted experience.

Choose a smaller model when low latency, modest hardware, or simple deployment is the priority. Choose a current hosted model when you need confirmed availability, predictable service-level behavior, or current pricing documentation. Within MiniMax's own lineup, newer catalog entries such as M3 and M2.7 may be more relevant for current hosted use, but the supplied research does not provide a detailed feature-by-feature comparison. For visual, audio, video, or music work, use a model designed for those modalities rather than expecting M2 to provide them natively.

Pricing and availability

The launch API price was $0.30 per million input tokens and $1.20 per million output tokens, with free access temporarily offered at launch. These figures should be treated as historical launch pricing, not a confirmed current rate. The current hosted lifecycle of the original M2 is unverified in the supplied research.

The open-weight model is available through MiniMax's official GitHub and Hugging Face repositories under a modified MIT license. Before deployment, verify the repository version, serving instructions, license conditions, hardware requirements, and whether the desired inference framework preserves the model's reasoning and tool-calling conventions.


Answers to Frequently Asked Questions

How much did MiniMax M2 API access cost at launch?
The launch API pricing was $0.30 per million input tokens and $1.20 per million output tokens, with temporary free access offered at launch. These figures are historical and should not be treated as current pricing; hosted availability and rates should be verified before deployment.
Can MiniMax M2 be deployed locally?
Yes. MiniMax M2 is available as an open-weight model through the official MiniMax-AI and Hugging Face repositories, with deployment guidance referencing frameworks such as vLLM, SGLang, and MLX. However, its approximately 230 billion total parameters make local deployment demanding and generally require substantial GPU memory or multi-GPU infrastructure.
How large is MiniMax M2's context window?
MiniMax M2 has a verified maximum sequence length of 196,608 tokens. This supports large codebases, extensive documentation, and long agent histories, although applications should still use sensible context management and retrieval strategies.
Is MiniMax M2 a multimodal model?
No. MiniMax M2 is text-only in the supplied specification: it accepts text input and produces text output. Native image, audio, video, music, speech, and embedding capabilities are not verified for this model.
What is MiniMax M2 designed for?
MiniMax M2 is an open-weight mixture-of-experts language model designed primarily for coding, reasoning, tool use, and long-running agent workflows. It is intended for tasks such as code generation, code repair, repository analysis, terminal operations, browser interactions, and multi-step planning.


Sources 6
Provider

About MiniMax