What is MiniMax M2?
MiniMax M2 is an open-weight mixture-of-experts language model released by MiniMax on October 27, 2025. It is designed chiefly for coding, reasoning, tool use, and agent workflows rather than for image, audio, or video generation.
The model has approximately 230 billion total parameters, but activates about 10 billion parameters for each token. This mixture-of-experts design can reduce the computation needed for each individual token compared with a dense model of the same total size. It does not, however, make the model small: local deployment still requires substantial GPU memory and generally multi-GPU infrastructure.
MiniMax M2 is distributed as an open-weight model through the official MiniMax-AI repository and Hugging Face model repository. The model is available under a modified MIT license, according to the supplied official materials.
Primary purpose and position in the MiniMax lineup
M2 occupies a specialist position in MiniMax's model portfolio. Its design emphasizes software engineering, tool-using agents, and long-running tasks that involve repeated planning and execution. It is not presented here as MiniMax's general consumer product or as a multimodal content-generation system.
MiniMax's current public catalog emphasizes newer models such as M3 and M2.7, so M2 should be treated as a legacy or independently useful open-weight model rather than assumed to be the provider's current flagship hosted option. Its open weights remain relevant for developers who want to inspect, adapt, or self-host the model, but its current hosted availability and pricing should be verified before building a production integration.
Core capabilities
M2 generates text and supports reasoning-oriented workflows. Its most relevant capabilities include:
- Coding: software development, code repair, multi-file changes, and codebase analysis.
- Reasoning: multi-step planning and interleaved thinking for tasks that require several intermediate decisions.
- Tool use: workflows involving shell commands, browser interactions, Python execution, retrieval, and MCP tools.
- Long-context processing: analysis of large codebases, documentation sets, and extended task histories.
- Agent operation: repeated cycles of planning, tool calls, observation, and further action.
Interleaved thinking means that reasoning is integrated between ordinary responses and actions rather than treated as a single isolated step. MiniMax's materials indicate that thinking content represented with think tags should be retained in the conversational history when M2 is used locally or through compatible APIs. Implementations therefore need to follow the model's documented prompting and history conventions instead of silently discarding those sections.
Context window and output limits
The verified maximum sequence length for MiniMax M2 is 196,608 tokens. A token is a piece of text used by the model, so this limit includes the relevant input and conversation history as well as generated content, depending on the serving implementation.
This unusually large context is useful for tasks such as reviewing a substantial repository, keeping a long agent trace available, or combining source files with technical documentation. It does not guarantee that every detail in a very large prompt will receive equal attention, and it does not remove the need for sensible retrieval and context management.
A separate maximum output-token limit was not clearly documented in the supplied first-party materials. Developers should therefore avoid assuming that the full context window can be emitted as one response and should confirm the limit exposed by the selected serving framework or hosted endpoint.
Coding and agent workflows
M2 is most distinctive when it is connected to tools. The documented examples include terminal or shell operations, browser actions, Python execution, retrieval, and MCP-based tools. These capabilities allow an application to delegate more than code completion: the model can help inspect files, plan changes, run commands, examine results, and continue working through a multi-step task.
In practice, a coding agent built around M2 might use the model to:
- Inspect a repository and identify the files relevant to a bug.
- Propose a repair plan and request or execute tool actions.
- Modify several files while preserving the broader project structure.
- Run tests or scripts through a shell or Python tool.
- Interpret failures and revise the implementation.
Tool access does not mean the model independently guarantees safe execution. Shell commands, browser access, file changes, and code execution should be isolated and permissioned by the application. The model's ability to call tools is a capability of the model-and-harness combination, not proof that every deployment includes the same tools or safety controls.
Supported modalities and output types
MiniMax M2 is text-only in the supplied specification. It accepts text input and produces text output. The research does not verify native image, audio, or video input, and it does not describe native image, audio, video, music, speech, or embedding output for this model.
This makes M2 a poor fit when the central task is visual understanding, media generation, voice interaction, or multimodal content creation. It can still participate in workflows that call external tools for those tasks, but that is different from having those modalities built into the model itself.
Speed, cost, and deployment trade-offs
The approximately 10-billion active-parameter figure helps explain why M2 can be attractive for inference efficiency relative to a dense 230-billion-parameter model. The total model size remains large, however, so the active-parameter count should not be interpreted as a low hardware requirement.
At launch, MiniMax listed API pricing of $0.30 per million input tokens and $1.20 per million output tokens. These were launch prices, and the original hosted M2 pricing is not confirmed as current. MiniMax's newer catalog pages emphasize other models, so prospective users should verify whether M2 is still hosted, under what endpoint, and at what rate before estimating operating costs.
For local deployment, the trade-off is different. Open weights can provide control over hosting and experimentation, but the model's total size makes infrastructure more demanding than the active-parameter figure alone suggests. Official deployment guidance references frameworks including vLLM, SGLang, and MLX. Hardware requirements will depend on quantization, parallelism, context size, and serving configuration; the supplied research does not establish one universal memory requirement.
Strengths and limitations
Where M2 is strong
- Specialization for coding agents: its intended workloads include code generation, code repair, tool use, and multi-step software tasks.
- Large context: a 196,608-token maximum sequence length supports extensive code and task histories.
- Open-weight access: developers can experiment with local deployment rather than relying exclusively on a hosted endpoint.
- Tool-oriented reasoning: the model is designed for workflows that combine planning with shell, browser, Python, retrieval, and MCP actions.
- Potential token-efficiency: the MoE architecture activates approximately 10 billion parameters per token rather than all 230 billion.
Where M2 is limited
- Large deployment footprint: 230 billion total parameters make self-hosting demanding despite the lower active count.
- Text-only operation: it is not a native image, audio, video, or music model.
- Unclear current hosted status: present API availability and pricing are not confirmed by the supplied current catalog information.
- Incomplete documented limits: the maximum output-token limit, knowledge cutoff, structured-output support, caching, batch API support, and fine-tuning status are not clearly established.
- Operational safety requirements: applications must control permissions and validate tool actions rather than treating model-generated commands as automatically safe.
Best use cases
MiniMax M2 is a strong candidate for developers evaluating open-weight models for:
- Multi-file software development and code repair.
- Long-context codebase or documentation analysis.
- Terminal, browser, retrieval, and code-execution agents.
- Research assistants that need repeated tool calls and intermediate planning.
- Self-hosted experiments with a large mixture-of-experts coding model.
- Applications where open weights and deployment control matter more than a simple managed API.
It is less appropriate for a lightweight local assistant, a media-generation application, or a production system that requires a clearly documented current endpoint, output limit, structured-output guarantee, or lifecycle commitment.
When to choose MiniMax M2
Choose M2 when the central problem is text-based coding or agent automation and you can accommodate the infrastructure needed for a large open-weight model. It is especially appealing when a long context, tool-driven workflow, and the ability to experiment with model weights are more important than a turnkey hosted experience.
Choose a smaller model when low latency, modest hardware, or simple deployment is the priority. Choose a current hosted model when you need confirmed availability, predictable service-level behavior, or current pricing documentation. Within MiniMax's own lineup, newer catalog entries such as M3 and M2.7 may be more relevant for current hosted use, but the supplied research does not provide a detailed feature-by-feature comparison. For visual, audio, video, or music work, use a model designed for those modalities rather than expecting M2 to provide them natively.
Pricing and availability
The launch API price was $0.30 per million input tokens and $1.20 per million output tokens, with free access temporarily offered at launch. These figures should be treated as historical launch pricing, not a confirmed current rate. The current hosted lifecycle of the original M2 is unverified in the supplied research.
The open-weight model is available through MiniMax's official GitHub and Hugging Face repositories under a modified MIT license. Before deployment, verify the repository version, serving instructions, license conditions, hardware requirements, and whether the desired inference framework preserves the model's reasoning and tool-calling conventions.

