MiniMax M2

MiniMax M2.7

by MiniMax · Current; available through MiniMax API, MiniMax Agent, and downloadable open weights

MiniMax M2.7 is an open-weight mixture-of-experts model focused on repository-scale coding, tool-using agents, long-context reasoning, debugging, and professional productivity. It offers a 196K-token context window, downloadable weights, API access, and standard pricing of $0.30 per million input tokens and $1.20 per million output tokens. Its main trade-offs are demanding local hardware requirements and text-only output.

Text Reasoning Coding
MiniMax M2.7 is a 230-billion-parameter mixture-of-experts language model with approximately 10 billion active parameters and a 196K-token context window. Released on March 18, 2026, it targets repository-scale coding, production debugging, professional productivity, tool-driven automation, and long-running agent workflows. It is available through the MiniMax API, MiniMax Agent, and downloadable model weights.
Outputs

What MiniMax M2.7 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiniMax M2
Model type Coding
Context window 197K tokens
Release date 2026-03-18
Status Current; available through MiniMax API, MiniMax Agent, and downloadable open weights
Knowledge cutoff notes

No authoritative knowledge-cutoff date for the exact MiniMax M2.7 model was identified in the reviewed first-party documentation.

Model notes

MiniMax M2.7 is a mixture-of-experts model with 230 billion total parameters and approximately 10 billion active parameters. The model supports tool calling and is available through the MiniMax API, MiniMax Agent, and downloadable weights. MiniMax also offers MiniMax-M2.7-highspeed, described as a faster variant with identical results. The standard API model identifier is MiniMax-M2.7. The 196K context value is documented in the official deployment ecosystem; an exact maximum output-token limit was not found in the authoritative sources reviewed. The editorial scores are comparative estimates, not provider-published ratings.

Cost

Model pricing

Input $0.30 per 1 million tokens; cache read $0.06 per 1 million tokens; cache write $0.375 per 1 million tokens
Output $1.20 per 1 million tokens
Model guide

MiniMax M2.7: An Open-Weight Model for Long-Running Coding Agents

MiniMax M2.7 is an open-weight, mixture-of-experts language model designed for software engineering, complex tool use, office productivity workflows, and long-running autonomous agents. It combines a 196K-token context window with tool calling, downloadable weights, API access, and relatively low token pricing, but its large 230-billion-parameter architecture makes local deployment demanding and it does not provide native image, audio, video, or music generation.

What is MiniMax M2.7?

MiniMax M2.7 is an agent-focused language model from MiniMax. Its primary role is not image creation or conversational entertainment; it is designed to carry out substantial software engineering and professional work with the help of tools, files, and multi-step instructions.

The model is a mixture-of-experts, or MoE, system. It has 230 billion total parameters, while approximately 10 billion parameters are active for a given operation. In practical terms, this architecture gives M2.7 a very large overall model capacity without activating every parameter for every token. That does not make local deployment lightweight, however: the published weights require approximately 220 GB of memory before accounting for runtime overhead and the key-value cache used to track long conversations.

MiniMax released M2.7 on March 18, 2026. The model is available through the MiniMax API Platform and MiniMax Agent, and MiniMax also provides downloadable weights through the official MiniMaxAI/MiniMax-M2.7 repository. The standard API model identifier is MiniMax-M2.7.

Where M2.7 fits in MiniMax's lineup

M2.7 belongs to MiniMax's text-model family and is positioned around coding, agentic execution, and professional productivity. It is separate from MiniMax products focused on video, image, speech, music, or general consumer interaction. The model can be used inside MiniMax Agent, but the underlying capability is still text generation with tool support rather than direct media generation.

MiniMax also offers MiniMax-M2.7-highspeed. According to the supplied provider information, the high-speed variant is intended for higher throughput and is described as producing identical results to the standard version. It is therefore a speed-and-price option within the same M2.7 offering rather than a separate capability tier.

Core capabilities and reported performance

M2.7 is intended for tasks where a model must understand a large body of information, plan several steps, use tools, and revise its work. MiniMax specifically positions it for software engineering, production debugging, code security, machine-learning workflows, Android development, log analysis, and full-project delivery.

Its coding focus includes repository-level work rather than only short code completion. Suitable assignments can include tracing a problem across multiple files, modifying an existing project, running checks through an agent harness, and iterating on the result. The model also supports complex tool-oriented workflows, including Agent Teams, dynamic tool search, and persistent-memory agent harnesses, according to MiniMax's published materials.

MiniMax reports a 56.22% score on SWE-Pro, 55.6% on VIBE-Pro, and 57.0% on Terminal Bench 2. These are provider-reported benchmark figures, not independent editorial ratings, and benchmark results should not be treated as a guarantee for every codebase or production environment.

For professional and office work, MiniMax reports a GDPval-AA ELO score of 1495. The provider describes support for multi-round editing of Word, Excel, and PowerPoint deliverables. That makes M2.7 relevant to workflows in which an agent must create a document, inspect feedback, and revise the output rather than produce a single response.

Context window, inputs, and outputs

The documented maximum context length is 196,608 tokens, commonly described as a 196K-token context window. A context window is the amount of text and related conversation state the model can consider in one request. This is useful for large repositories, long logs, technical documentation, extended task histories, and multi-step agent sessions.

The supplied documentation does not establish a separate maximum output-token limit for the exact M2.7 model, so no maximum output figure should be assumed. The model accepts text input and produces text output. The supplied specifications do not identify image, audio, or video input support for M2.7, and it is not a native image, audio, video, speech, or music generation model.

The large context window should also be distinguished from available memory. A long context requires additional key-value-cache memory during inference. Consequently, an application may be able to send a large request through an API without being able to run the same request efficiently on a local machine.

Reasoning, coding, and tool use

M2.7 is best understood as a reasoning-oriented work model rather than a model intended only to answer short factual questions. Its value comes from decomposing tasks, maintaining state across steps, interpreting tool results, and revising an implementation or document.

Tool calling is supported. This allows an application or agent harness to expose functions such as repository operations, test execution, file handling, or other approved actions. The model can then request a tool call in the course of completing a task. Tool calling does not mean that every deployment automatically has browser access or unrestricted system access: the application must provide the tools and enforce permissions.

MiniMax documents deployment with frameworks including vLLM, SGLang, and Transformers. The official tool-calling materials and downloadable weights make M2.7 relevant to teams that want more control over deployment than a hosted chat interface provides. At the same time, the approximately 220 GB weight requirement places it in a substantially different hardware category from smaller dense models.

The reviewed documentation does not establish a separate native web-search capability, a fine-tuning service, a batch API, or a distinct JSON mode for M2.7. Structured tool workflows may still be possible through an integration, but those should not be presented as verified native features unless the specific implementation documents them.

API access and deployment choices

The documented API endpoint is https://api.minimax.io/v1/text/chatcompletion_v2. Hosted API access is the most straightforward route for users who want M2.7's long context and coding capabilities without acquiring hardware capable of holding the model weights.

The standard model is priced for token usage, while the high-speed variant trades a higher token rate for higher throughput. Self-hosting offers greater control over infrastructure and data handling, but it requires substantially more memory and operational expertise. The choice is therefore not simply between two interfaces: it is a trade-off between hosted convenience and local control, with the hardware burden being a central consideration for local deployment.

MiniMax M2.7 pricing

Current MiniMax API pricing supplied for the standard M2.7 model is:

Usage typeStandard M2.7M2.7-highspeed
Input tokens$0.30 per 1 million tokens$0.60 per 1 million tokens
Output tokens$1.20 per 1 million tokens$2.40 per 1 million tokens
Cache reads$0.06 per 1 million tokens$0.06 per 1 million tokens
Cache writes$0.375 per 1 million tokens$0.375 per 1 million tokens

These are usage prices rather than monthly subscription prices. The high-speed option costs twice as much for input and output tokens in the supplied pricing, while its cache prices are listed as the same as the standard variant. For workloads that can tolerate lower throughput, standard M2.7 is the more economical choice. For latency-sensitive or high-volume agent systems, the high-speed variant may be preferable if its throughput advantage justifies the additional token cost.

Main strengths

  • Long context: The 196K-token window is suited to large repositories, long logs, extensive documentation, and persistent task histories.
  • Software engineering focus: The model is explicitly designed for repository work, debugging, security analysis, machine-learning workflows, Android development, and project delivery.
  • Agent and tool support: Tool calling, dynamic tool search, Agent Teams, and persistent-memory agent harnesses support multi-step workflows.
  • Open-weight availability: Downloadable weights provide an option for organizations that need more deployment control than a hosted API offers.
  • Low stated token rates: The standard API price is relatively economical for workloads that can use the regular-speed endpoint, especially when cache reads reduce repeated-context costs.
  • Professional document workflows: Provider-reported support for iterative Word, Excel, and PowerPoint work broadens its use beyond source code.

Limitations and risks to consider

  • Demanding self-hosting requirements: Approximately 220 GB is required for the weights, before runtime memory and long-context cache requirements. This rules out many ordinary developer machines.
  • Text-only model output: M2.7 is not the appropriate choice for directly generating images, videos, audio, speech, or music.
  • No verified knowledge cutoff: The reviewed authoritative material does not provide a knowledge-cutoff date for this exact model, so time-sensitive factual work should use current data sources or tools where appropriate.
  • Unverified additional API features: The supplied documentation does not establish native web search, fine-tuning, batch processing, or a separate JSON mode.
  • Benchmark interpretation: MiniMax's reported scores are useful signals about intended performance, but they are provider claims and do not guarantee results on a particular repository, language, test suite, or office workflow.
  • Agent risk: Tool-enabled automation can make consequential changes. Production deployments should limit permissions, inspect tool calls, and add human approval for destructive operations.

When to choose MiniMax M2.7

Choose M2.7 when the central problem is a long-running coding or productivity task that benefits from a large context, tool calls, and iterative execution. It is a strong candidate for repository-scale code changes, production debugging, log investigation, code-security workflows, document automation, and agents that need to preserve substantial task context.

The standard API version is the sensible starting point when cost matters more than maximum throughput. M2.7-highspeed is more appropriate when response speed or concurrent agent capacity is important and the higher input and output prices fit the budget. Hosted access is generally more practical than local deployment unless an organization already has infrastructure capable of holding a model of this size.

Another type of model may be more appropriate when the task requires direct media generation, a small local footprint, a documented native web-search feature, a verified knowledge cutoff, or a specifically documented structured-output and batch-processing interface. A smaller coding model may also be preferable for short edits, autocomplete, or lightweight automation where M2.7's long context and agentic capacity are unnecessary.

Bottom line

MiniMax M2.7 is a specialized open-weight language model for coding agents, tool-driven work, and long professional workflows. Its defining combination is a 196K-token context window, approximately 10 billion active parameters within a 230-billion-parameter MoE architecture, tool support, downloadable weights, and relatively low standard API token pricing. Those advantages come with clear boundaries: local operation requires substantial hardware, the model is text-only, and several commonly requested features are not verified in the supplied documentation. For users who need repository-scale reasoning and agent execution rather than media generation or lightweight chat, M2.7 is a focused option worth evaluating.


Answers to Frequently Asked Questions

Does MiniMax M2.7 support tool calling and media generation?
MiniMax M2.7 supports tool calling, allowing applications to provide functions for repository operations, testing, file handling, and other approved actions. It is a text-only model and is not designed to directly generate images, video, audio, speech, or music. Browser access and system access are not automatic; they must be provided and controlled by the host application.
How much does MiniMax M2.7 API access cost?
Standard MiniMax M2.7 pricing is $0.30 per 1 million input tokens and $1.20 per 1 million output tokens. Cache reads cost $0.06 per 1 million tokens, while cache writes cost $0.375 per 1 million tokens. The M2.7-highspeed variant costs $0.60 per million input tokens and $2.40 per million output tokens, with the same listed cache prices.
Can MiniMax M2.7 be run locally?
Yes, MiniMax provides downloadable weights and documents deployment with frameworks such as vLLM, SGLang, and Transformers. However, the weights require approximately 220 GB of memory before accounting for runtime overhead and long-context key-value cache requirements, so local deployment requires substantial hardware.
What is MiniMax M2.7 designed for?
MiniMax M2.7 is an agent-focused, open-weight language model designed for long-running software engineering and professional productivity tasks. It supports repository-level coding, debugging, code security, log analysis, machine-learning workflows, document automation, and other tool-driven, multi-step tasks.
How large is MiniMax M2.7's context window?
MiniMax M2.7 has a documented maximum context window of 196,608 tokens, commonly described as 196K tokens. This supports large code repositories, lengthy logs, extensive documentation, and persistent multi-step agent sessions.


Sources 7
Provider

About MiniMax