MiniMax M3

MiniMax M3

by MiniMax · Active; open-weight and available through the MiniMax API

Open-weight MiniMax model for coding agents, long-context reasoning, multimodal analysis, and tool-using workflows. M3 supports up to 1 million tokens of context, accepts text, images, and video, returns text, and is available through the MiniMax API or local deployment.

Text Reasoning Coding
MiniMax M3 is MiniMax’s flagship coding and agentic model. Its main distinction is the combination of a 1-million-token context window, native image and video understanding, controllable reasoning, tool invocation, and open-weight availability. It is intended for software-engineering agents, large repositories, long technical documents, multimodal analysis, and private deployments rather than native media generation.
Outputs

What MiniMax M3 can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Web search Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiniMax M3
Model type Multimodal
Context window 1M tokens
Release date 2026-06-01
Status Active; open-weight and available through the MiniMax API
Knowledge cutoff notes

No authoritative MiniMax source reviewed publishes an exact knowledge cutoff for MiniMax M3.

Model notes

MiniMax describes M3 as a natively multimodal coding and agentic model powered by MiniMax Sparse Attention. The official model page states that the API supports up to 1M tokens with a guaranteed minimum of 512K. The Hugging Face checkpoint is approximately 427B total parameters with approximately 23B active parameters and is distributed under the MiniMax Community License. The model supports controllable thinking modes including enabled, adaptive, and disabled. Reasoning, coding, speed, and cost scores are editorial comparative estimates rather than provider-issued ratings. The exact knowledge cutoff and maximum output-token limit were not verified from an authoritative first-party source.

Cost

Model pricing

Input $0.30 per million tokens for context up to 512K; $0.60 per million tokens for context from 512K to 1M. Listed rates are subject to MiniMax's current pricing terms.
Output $1.20 per million tokens for context up to 512K; $2.40 per million tokens for context from 512K to 1M. Listed rates are subject to MiniMax's current pricing terms.
Model guide

MiniMax M3: An Open-Weight Model for Long-Context Coding Agents

MiniMax M3 is an active, open-weight multimodal model from MiniMax designed for coding, long-context analysis, agentic workflows, and tool use. It accepts text, images, and video, supports up to 1 million tokens of context, returns text, and is available through the MiniMax API or downloadable weights for local deployment.

What is MiniMax M3?

MiniMax M3 is an active open-weight model released by MiniMax on June 1, 2026. It is built primarily for coding, autonomous or semi-autonomous agent workflows, long-context reasoning, and professional knowledge work. The model is available through the MiniMax Open Platform API and as downloadable weights through Hugging Face.

For a general user, the most important distinction is that M3 is an understanding and generation model whose output is text. It can analyze text, images, and video, reason through complex tasks, write or modify code, and interact with tools, but it does not directly generate images, video, audio, speech, or music.

Where M3 fits in MiniMax’s lineup

M3 is positioned as MiniMax’s coding- and agent-focused model rather than as a general consumer media-generation product. MiniMax’s broader ecosystem includes separate services for video, audio, image creation, and agentic work. M3’s role is narrower and more technical: it provides a long-context language and reasoning core for software development, document analysis, multimodal understanding, and tool-using systems.

That positioning matters when choosing the model. M3 may be a strong fit when the task requires reading a large amount of material, maintaining context across many steps, or operating through external tools. It is not the appropriate choice when the required result is a generated image, video clip, voice recording, or music track.

Key specifications

SpecificationMiniMax M3
ProviderMiniMax
Release dateJune 1, 2026
StatusActive; open-weight and available through the MiniMax API
Context lengthUp to 1,000,000 tokens
Input typesText, images, and video
Output typeText
Tool useSupported
Model sizeApproximately 427 billion total parameters, with approximately 23 billion active parameters in the published checkpoint
Maximum output tokensNot verified in the supplied authoritative sources

MiniMax describes M3 as natively multimodal, meaning image and video understanding are part of the model’s training and design rather than being presented only through an unrelated external wrapper. The official model information states that the API supports up to 1 million tokens, with a guaranteed minimum context of 512,000 tokens.

Long-context processing and architecture

M3 uses MiniMax Sparse Attention, or MSA, to handle very long inputs while reducing the computational work required compared with treating every token as equally connected to every other token. In practical terms, the large context window can help the model work with extensive codebases, lengthy technical documentation, large collections of records, or multi-step task histories without requiring the user to divide the material into many unrelated requests.

A 1-million-token context window is a capacity limit, not a guarantee that every long prompt will produce equally accurate results. Retrieval quality, prompt organization, irrelevant material, and the complexity of the requested task still affect the answer. For best results, users should identify the relevant files or sections clearly and ask for a concrete operation, such as locating a defect, proposing a refactor, or comparing implementation choices.

Reasoning, coding, and agent workflows

M3 supports controllable thinking modes. According to the supplied research, thinking can be enabled, disabled, or selected adaptively. This gives developers a way to trade off deliberation against response speed and computational cost: more deliberate reasoning may be useful for architecture changes or multi-step debugging, while a lighter mode may be sufficient for routine transformations or short answers.

The model is designed for tool invocation and multi-step execution. A tool-using application can provide functions for actions such as reading files, running commands, querying services, or interacting with other systems. M3 can then help decide which action to take, interpret the result, and continue the workflow. The model’s support for browsing-oriented and computer-use scenarios makes it relevant to coding agents and office automation, although the quality and safety of an agent also depend on the surrounding application, permissions, and confirmation controls.

For coding, M3 is suited to tasks such as navigating a large repository, explaining unfamiliar modules, generating implementation plans, writing code, reviewing changes, and working through multi-step fixes. Its long context is particularly relevant when the task depends on relationships between many files rather than on a short isolated code snippet.

Supported modalities: broad input, text-only output

M3 accepts text, images, and video as inputs. This allows use cases such as examining a screenshot of an interface, analyzing a diagram alongside written requirements, or reviewing visual material together with technical documentation. Video input can support analysis of events or demonstrations when the application supplies the video in a supported format.

The output side is more limited. M3 returns text and does not natively produce images, video, audio, speech, or music. An application can potentially connect its textual instructions to separate generation tools, but that would be a workflow combining M3 with other services rather than native M3 output. This distinction is important for teams looking for one model to perform both analysis and media generation.

API pricing and access options

MiniMax lists two API price bands based on context length. For requests with context up to 512,000 tokens, the listed price is $0.30 per million input tokens and $1.20 per million output tokens. For requests with context between 512,000 and 1 million tokens, the listed price rises to $0.60 per million input tokens and $2.40 per million output tokens. Cache-read pricing is listed separately by MiniMax.

Context rangeInputOutput
Up to 512K tokens$0.30 per million tokens$1.20 per million tokens
512K to 1M tokens$0.60 per million tokens$2.40 per million tokens

These are usage prices rather than a fixed consumer subscription price, and users should check MiniMax’s current pricing terms before deployment. The higher rate for very large contexts creates a practical trade-off: M3 can reduce the need to split large tasks into many requests, but sending close to the maximum context can cost more than using shorter, carefully selected inputs.

Developers can also download the published weights for local or private deployment. The checkpoint is distributed under the MiniMax Community License. Local deployment can provide greater control over data handling and infrastructure, but the approximately 427-billion-parameter checkpoint, with approximately 23 billion active parameters, requires substantial hardware and operational expertise. The active-parameter figure should not be interpreted as meaning that the full checkpoint is small or easy to run.

Main strengths and trade-offs

  • Very large context: The stated 1-million-token capacity is useful for large repositories, long documents, and extended agent histories.
  • Multimodal understanding: Text, image, and video inputs support analysis beyond ordinary text-only coding prompts.
  • Agent orientation: Tool invocation, controllable thinking, and multi-step execution are central to the model’s intended use.
  • Open-weight availability: Downloadable weights offer an option for organizations that need local or private deployment.
  • Usage-based cost: The listed token prices are relatively low for many requests, but the 512K-to-1M context tier costs more and large prompts can accumulate significant usage.
  • Infrastructure burden: Local deployment requires substantial resources because of the size of the published checkpoint.
  • Text-only generation: M3 cannot replace a dedicated image, video, audio, or music generation model.

Editorial assessments in the supplied research rate M3 highly for reasoning and coding, with strong scores for speed and cost as comparative estimates. Those ratings are not MiniMax-published benchmarks or guarantees. Actual latency and expense will depend on prompt length, reasoning mode, context tier, API conditions, deployment hardware, and the complexity of the tool workflow.

Best use cases for MiniMax M3

  • Long-running software-engineering agents that inspect, edit, and reason about multiple files.
  • Large repositories where important context is distributed across many modules.
  • Analysis of long technical documents, specifications, or collections of project material.
  • Multimodal document work involving text, screenshots, diagrams, or video.
  • Tool-using automation that needs planning, function calls, and follow-up interpretation.
  • Private or controlled deployments where downloadable weights are more suitable than sending all data to a hosted service.

A practical coding request might ask M3 to inspect a repository, identify where a data-flow problem originates, explain the relevant files, propose a patch, and then use available tools to test the change. A document workflow might provide a large specification and supporting diagrams, then ask for a traceability matrix or an implementation checklist.

When to choose MiniMax M3

Choose M3 when the task benefits from a large working memory, multimodal understanding, coding ability, and tool-based execution in one model. It is especially compelling when splitting the task into many smaller prompts would lose important relationships or create substantial orchestration work.

Choose a smaller or more specialized model when the task is a simple classification, short rewrite, routine extraction, or high-volume operation where maximum context and extended reasoning are unnecessary. A shorter prompt and lighter reasoning mode may also be preferable when response speed and predictable cost matter more than complex planning.

Choose a dedicated media-generation model when the desired result is an image, video, voice, speech, or music file. M3 can analyze those media types in supported inputs, but its native output remains text. For local deployment, choose M3 only if the organization can support the hardware, serving stack, license requirements, and operational maintenance associated with a very large checkpoint.

Limitations and unverified details

The exact maximum output-token limit and knowledge cutoff for M3 were not verified in the authoritative sources supplied for this page. They should not be assumed from the 1-million-token context figure. Context length describes the total working window, while maximum output is a separate limit.

The model’s tool-use capability also does not mean that it can safely perform unrestricted actions by itself. Developers remain responsible for defining available functions, controlling permissions, validating arguments, protecting sensitive data, and requiring confirmation for consequential operations. Likewise, open weights provide deployment flexibility but do not remove the need for evaluation, monitoring, and appropriate security controls.

Overall, MiniMax M3 is best understood as a long-context, multimodal-understanding model for coding and agents. Its strongest practical case is a workflow that needs to read a great deal, reason across multiple steps, use tools, and return a detailed textual result. Its main compromises are the cost of very large contexts, the infrastructure demands of local deployment, unverified output and knowledge-cutoff details, and the absence of native non-text generation.


Answers to Frequently Asked Questions

Can MiniMax M3 be deployed locally?
Yes. MiniMax M3's published weights can be downloaded from Hugging Face under the MiniMax Community License. However, the checkpoint contains approximately 427 billion total parameters, including about 23 billion active parameters, so local deployment requires substantial hardware, infrastructure, and operational expertise.
How much does MiniMax M3 API access cost?
For contexts up to 512,000 tokens, MiniMax lists pricing of $0.30 per million input tokens and $1.20 per million output tokens. For contexts between 512,000 and 1 million tokens, the listed prices are $0.60 per million input tokens and $2.40 per million output tokens. Cache-read pricing is listed separately, and users should verify current pricing before deployment.
Can MiniMax M3 generate images, video, audio, or music?
No. MiniMax M3 accepts text, images, and video as inputs, but its native output is text only. Generating images, video, audio, speech, or music requires connecting M3 to separate specialized services.
What is MiniMax M3 designed for?
MiniMax M3 is an open-weight model designed primarily for coding, long-context reasoning, autonomous or semi-autonomous agents, multimodal understanding, and professional knowledge work. It can analyze text, images, and video, use tools, and generate text-based responses.
How large is MiniMax M3's context window?
MiniMax M3 supports a context window of up to 1 million tokens, with a guaranteed minimum context of 512,000 tokens according to the supplied model information. This capacity is useful for large codebases, lengthy documents, and extended agent workflows, but it does not guarantee perfect accuracy on every long prompt.


Sources 5
Provider

About MiniMax