MiniMax M2

MiniMax M2.1

by MiniMax · Legacy but currently available

MiniMax M2.1 is an open-weight Mixture-of-Experts model for multilingual programming, coding agents, tool use, long-context workflows, and office automation. It remains API-accessible as a legacy model at $0.30 per million input tokens and $1.20 per million output tokens, but it does not support native image, audio, or video tasks.

Text Reasoning Coding
MiniMax M2.1 is a coding- and agent-focused language model from MiniMax. It is intended for software development, long-running tool-using tasks, technical writing, and workplace automation rather than image, audio, or video generation. The model is available through MiniMax's Open Platform and as downloadable weights for local deployment. Although MiniMax now labels M2.1 as a legacy model, its API remains accessible and its relatively low token prices make it relevant for cost-sensitive coding and automation workloads.
Outputs

What MiniMax M2.1 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiniMax M2
Model type Coding
Context window 205K tokens
Release date 2025-12-23
Status Legacy but currently available
Knowledge cutoff notes

MiniMax has not published a direct knowledge-cutoff date for the exact M2.1 model in the authoritative documentation reviewed. Web search or server-side retrieval, where separately enabled, does not change the underlying model knowledge cutoff.

Model notes

MiniMax M2.1 is an open-weight Mixture-of-Experts model with approximately 230 billion total parameters and approximately 10 billion activated parameters. MiniMax emphasizes multilingual coding, web and app development, interleaved thinking, long-horizon tool use, and office automation. The current MiniMax API documentation lists M2.1 as available through Anthropic-compatible and OpenAI-compatible interfaces with a 204,800-token context window and approximately 60 tokens per second output speed. The current pricing page classifies it as a legacy model but does not list a shutdown date. The downloadable Hugging Face configuration specifies a 196,608-token maximum position length for local weights, which differs from the provider's 204,800-token API context listing. Editorial scores are comparative estimates, not vendor specifications. The M2.1-highspeed variant is a separate faster model record and should not be merged with M2.1.

Cost

Model pricing

Input $0.30 per 1 million tokens; prompt-cache read $0.03 per 1 million tokens; prompt-cache write $0.375 per 1 million tokens
Output $1.20 per 1 million tokens
Model guide

MiniMax M2.1: An Open-Weight Coding Model for Agent Workflows

MiniMax M2.1 is an open-weight Mixture-of-Experts language model designed for multilingual software engineering, coding agents, tool use, application development, and office automation. Released on December 23, 2025, it remains available through the MiniMax Open Platform but is currently classified as a legacy model. Its API offers a 204,800-token context window and pricing of $0.30 per million input tokens plus $1.20 per million output tokens.

What is MiniMax M2.1?

MiniMax M2.1 is an open-weight Mixture-of-Experts, or MoE, language model released by MiniMax on December 23, 2025. In an MoE model, a large collection of learned parameters is available, but only a portion is activated for each request. MiniMax's published description identifies M2.1 as having approximately 230 billion total parameters and approximately 10 billion activated parameters. Those figures describe the model architecture; they do not by themselves guarantee a particular level of application performance.

The model is built primarily for software engineering and agentic work. That includes writing and reviewing code, modifying applications, following multi-step instructions, calling tools, and carrying out structured office tasks. It supports text-based reasoning and tool calls rather than producing images, audio, or video.

M2.1 improves on the earlier MiniMax M2 generation in areas MiniMax emphasizes as important for practical work: multilingual programming, web and mobile application development, long-horizon tool use, technical writing, and office-task automation. The model also uses interleaved thinking, allowing reasoning and tool-use steps to be combined during an agent workflow.

Current status and positioning

M2.1 remains available through the MiniMax Open Platform, including OpenAI-compatible and Anthropic-compatible interfaces. However, the current MiniMax pricing catalog classifies it as a legacy model. The supplied documentation does not provide a deprecation or shutdown date, so legacy status should be treated as a positioning and lifecycle warning rather than proof that access has ended.

For practical evaluation, this means M2.1 can still be a useful, inexpensive model for existing integrations and coding agents, but teams beginning a long-lived deployment should also check MiniMax's current model catalog and lifecycle notices. A newer model may be a better choice if future availability, feature continuity, or access to the provider's latest capabilities is more important than M2.1's current price and established integrations.

Core capabilities for coding and agents

M2.1 is optimized for programming in languages including Rust, Java, Go, C++, Kotlin, Objective-C, TypeScript, and JavaScript. The documented scope also includes web development, mobile application development, technical writing, and multi-step workplace tasks.

Its main practical distinction is not simply that it can generate code. It is designed to participate in a longer workflow: inspect a task, reason about the next step, call an available tool, interpret the result, and continue. Tool calling lets an application expose functions such as file operations, search, testing, or deployment actions to the model. The model produces the requested structured tool call, while the surrounding application actually executes it and returns the result.

M2.1 supports agent frameworks and coding tools including Claude Code, Cline, Kilo Code, Roo Code, Droid, and BlackBox. Compatibility with a tool does not mean that every integration has identical features or reliability; the surrounding framework, prompt, permissions, and tool definitions remain important parts of the system.

Reasoning and interleaved thinking

MiniMax positions M2.1 for complex instruction following and long-horizon tasks. Its interleaved-thinking behavior is useful when a request requires several decisions separated by tool results, such as updating a codebase, checking an output, and then correcting an implementation.

The available research does not provide a standardized benchmark result or a provider-published reasoning score for M2.1. An editorial comparison rates its reasoning at 8 out of 10 and coding at 9 out of 10, but those are comparative editorial judgments rather than MiniMax specifications. They should be used as directional assessments, not as guaranteed performance levels.

Context window and deployment options

The MiniMax API documentation lists a 204,800-token context window for M2.1. A context window is the amount of input and conversation material the service can consider in one request, including instructions, code, previous messages, tool results, and other supplied text. This large limit is relevant to repository analysis, lengthy technical documents, and agent sessions that accumulate substantial tool output.

The API documentation also gives an approximate output speed of 60 tokens per second. Actual throughput can vary with service conditions, request size, streaming behavior, and account or platform limits. The research does not provide a model-specific maximum output-token limit, so applications should not assume that the full context window can be returned as output.

MiniMax also publishes downloadable weights through Hugging Face for local deployment. Documentation covers inference systems including Transformers, vLLM, and SGLang. The local model configuration specifies a 196,608-token maximum position length, which differs from the provider's 204,800-token API context listing. These values should not be treated as interchangeable: the API limit describes MiniMax's hosted service, while the configuration value applies to the published local weights and their supported runtime.

Pricing for API use

The current pay-as-you-go catalog lists M2.1 as a legacy model at the following standard rates:

Usage typePrice
Input tokens$0.30 per million tokens
Output tokens$1.20 per million tokens
Prompt-cache reads$0.03 per million tokens
Prompt-cache writes$0.375 per million tokens

Input and output prices are different because generated tokens generally require more model computation than tokens supplied in the prompt. The low input rate can be attractive for applications that repeatedly analyze large codebases or documents, while the output rate matters more for workflows that generate extensive code, explanations, or tool arguments.

Prompt caching can reduce the cost of repeated context when the same material is reused, but the cache read and write charges should be included in cost calculations. MiniMax lists an M2.1-highspeed variant separately; it is a different model record and should not be assumed to have the same behavior or pricing.

Supported inputs and outputs

M2.1 is a text model with tool-use support. The documented interface accepts text and tool-call content through the Anthropic-compatible API. It does not support image or video input through that interface, and the model is not documented as an audio, speech, image, video, music, embedding, or moderation model.

  • Text input: Supported.
  • Text output: Supported.
  • Tool or function calls: Supported.
  • Image, audio, and video input: Not supported in the documented interface.
  • Image, audio, video, music, speech, embedding, or moderation output: Not supported.

This makes M2.1 suitable for a text-centered coding agent, but not for an application that needs native visual understanding, speech generation, image creation, or multimodal document analysis. Sending a screenshot or audio file to a surrounding application does not turn M2.1 into a multimodal model unless the provider explicitly supplies a compatible input path.

Main strengths and trade-offs

M2.1's strongest case is the combination of coding specialization, tool use, long context, open-weight availability, and relatively low API pricing. It can be considered when a team needs a model to work through substantial text-based project context or repeated agent steps without paying premium rates for every request.

  • Programming breadth: MiniMax specifically highlights several mainstream and systems languages, including Rust, Go, C++, Java, Kotlin, TypeScript, and JavaScript.
  • Agent workflow support: Tool calls and interleaved thinking are useful for coding assistants and multi-step automation.
  • Long-context API access: The documented 204,800-token window can accommodate large prompts, code collections, and tool histories.
  • Deployment flexibility: Hosted APIs are available alongside downloadable weights for local inference.
  • Cost: The standard input and output prices are low enough to support cost-sensitive automation, especially when prompt caching is effective.

There are also important trade-offs. The model is legacy in the current pricing catalog, so a newer model may offer a clearer long-term direction. Its approximately 60-token-per-second advertised speed is useful but not necessarily the fastest option for latency-sensitive applications. Local deployment also requires suitable infrastructure and an inference stack; publishing weights does not eliminate the operational work of hosting a large MoE model.

Limitations to consider

M2.1's modality limits are significant. It is not the right primary model for visual inspection, video understanding, audio processing, speech synthesis, music generation, or image generation. It also should not be selected when a documented maximum output-token limit or knowledge-cutoff date is a firm requirement: MiniMax has not published either detail for the exact model in the authoritative material reviewed.

The difference between the API context listing and the local configuration is another operational detail. Hosted applications should use the API's documented limit, while local deployments should follow the limits supported by the chosen weights and runtime. Exceeding practical memory or inference constraints may require reducing context even if the nominal position length is higher.

As with other coding models, generated code and tool actions require testing and permission controls. The supplied research does not establish a guaranteed reliability level or benchmark outcome. M2.1 can produce useful code and multi-step plans, but applications should validate outputs, restrict tools to the minimum required permissions, and handle failed or malformed tool calls.

When to choose MiniMax M2.1

Choose M2.1 when the main workload is text-based software engineering or automation and the following characteristics matter:

  • You need multilingual coding support across web, mobile, systems, or application development.
  • You want a model that can call tools and continue through multi-step agent tasks.
  • Your prompts or tool histories are large enough to benefit from a long context window.
  • Low token cost is more important than access to the newest model generation.
  • You want to evaluate downloadable weights or operate a local deployment using supported inference tools.

Another option may be more appropriate when the application requires native image, audio, or video capabilities; when the lowest possible latency is the priority; or when a current, non-legacy model with a clearer support horizon is required. A specialized multimodal model is a better fit for visual or audio tasks, while a newer coding model may be preferable for a greenfield system where lifecycle continuity matters more than M2.1's price.

Bottom line

MiniMax M2.1 is a focused text model for coding agents, software engineering, tool use, and long-running automation. Its 204,800-token hosted context, open-weight release, broad programming-language coverage, and low standard token prices make it a practical candidate for cost-conscious development workflows. Its legacy classification, lack of documented maximum output and knowledge-cutoff details, and absence of native multimodal support define the boundaries of that use case. It is best evaluated as an accessible coding and agent model rather than as a general-purpose model for every modality.


Answers to Frequently Asked Questions

When should you choose MiniMax M2.1 for an AI application?
MiniMax M2.1 is a suitable choice when an application needs cost-effective text-based coding, long-context repository analysis, multilingual programming, downloadable weights, or multi-step tool-using agents. A newer model may be preferable when long-term lifecycle support, the lowest latency, or native multimodal capabilities are more important.
Does MiniMax M2.1 support image, audio, or video input?
No. MiniMax M2.1 is documented as a text model with text output and tool-call support. Its documented interface does not support image, audio, or video input, and it is not intended for image, speech, music, video, embedding, or moderation generation.
How much does MiniMax M2.1 cost to use through the API?
The current standard pricing lists MiniMax M2.1 at $0.30 per million input tokens, $1.20 per million output tokens, $0.03 per million prompt-cache reads, and $0.375 per million prompt-cache writes. The model is currently classified as legacy in the pricing catalog.
What is MiniMax M2.1 designed for?
MiniMax M2.1 is an open-weight Mixture-of-Experts language model designed primarily for software engineering, coding assistants, tool use, and multi-step agent workflows. It supports tasks such as writing and reviewing code, modifying applications, technical writing, and structured office automation.
What is the context window of MiniMax M2.1?
The MiniMax API lists a 204,800-token context window for M2.1. The downloadable local configuration specifies a 196,608-token maximum position length, so hosted API limits and local deployment limits should be treated separately.


Sources 7
Provider

About MiniMax