Mistral Moderation

Mistral Moderation 2

by Mistral AI · Active; generally available; Premier

Mistral Moderation 2 is Mistral AI's active text moderation model for harmful-content and policy classification. It supports 128k-token inputs, multi-turn conversations, jailbreaking detection, batch processing, configurable guardrails, and free listed API pricing, but not media moderation or general text generation.

Reasoning Coding
Mistral Moderation 2 is a hosted safety-classification model for applications that need to assess text before or after it reaches an AI system or user. Released on March 1, 2026, it replaces the deprecated Mistral Moderation 2411 model and adds a larger context window, expanded policy categories, and detection of attempts to bypass safety policies. Its canonical API identifier is mistral-moderation-2603.
Inputs

What it can understand

Text
Capabilities

Supported features

Structured output Batch API
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family Mistral Moderation
Model type Moderation
Context window 131K tokens
Release date 2026-03-01
Status Active; generally available; Premier
Knowledge cutoff notes

Mistral does not publish a separate knowledge-cutoff date for this specialized moderation model. It is a classification service whose current behavior and policy categories may be updated over time.

Model notes

Canonical model ID is mistral-moderation-2603. The model classifies raw text and conversational content and returns policy-category scores and decisions. Documented categories include sexual, hate and discrimination, violence and threats, dangerous, criminal, self-harm, health, financial, law, PII, and jailbreaking. Mistral's custom guardrails use the moderation_llm_v2 configuration backed by this model. The moderation endpoint may be continually improved, so applications depending on category_scores may require recalibration. Mistral Moderation 2411 was deprecated on March 31, 2026.

Cost

Model pricing

Input Free
Output Free
Model guide

Mistral Moderation 2: Text Safety Classification and Jailbreak Detection

Mistral Moderation 2 is Mistral AI's current specialized moderation model for classifying harmful or policy-sensitive text and multi-turn conversations. With a 128k-token context window, jailbreaking detection, moderation and batch endpoints, and free listed input and output pricing, it is designed for content filtering and configurable safety guardrails rather than general-purpose text generation.

What Mistral Moderation 2 does

Mistral Moderation 2 is a specialized classification model from Mistral AI. It examines raw text or a sequence of conversational messages and returns moderation decisions and category scores. An application can use those results to allow content, send it for human review, apply a warning, or block it.

This is not a conversational model that writes answers for end users. It is a safety layer intended to evaluate content in systems such as chatbots, social platforms, support tools, and AI agents. The model's canonical identifier is mistral-moderation-2603, and Mistral exposes it through its Moderations API, classifiers interface, and batch processing endpoint.

Policy categories and jailbreaking detection

The model is documented for a broad set of policy categories. These include sexual content, hate and discrimination, violence and threats, dangerous content, criminal content, self-harm, health, financial topics, legal topics, personally identifiable information, and jailbreaking attempts.

Jailbreaking detection is one of the model's most relevant additions. In this context, a jailbreak is an attempt to manipulate an AI system into ignoring or bypassing its safety instructions, often through role-playing, adversarial wording, or carefully constructed prompts. Mistral Moderation 2 can provide a signal that an input is attempting this type of policy circumvention. That makes it useful not only for moderating user-generated content, but also for checking prompts submitted to another language model.

Category scores should be treated as risk signals rather than automatic statements of fact. Applications need to choose thresholds appropriate to their audience and consequences. A platform with strict child-safety requirements may prefer more human review, while a lower-risk internal workflow may tolerate a higher threshold before blocking content.

Context window and supported inputs

Mistral documents a 128k-token context window, equivalent to 131,072 tokens in the structured model record. This gives the model room to evaluate long documents or extended multi-turn conversations without reducing the input to only its latest message.

The supported inputs are text and conversational message lists. A raw-text moderation request can accept one string or a list of strings for small batched requests. A separate conversational endpoint can evaluate a sequence of user and assistant messages. This distinction is useful when the safety decision depends on context: a single sentence may appear harmless in isolation but become problematic when considered alongside earlier turns.

The supplied specifications do not identify image, audio, or video input support. Mistral Moderation 2 should therefore be treated as a text moderation model, not as a general media-moderation system. It also does not produce images, audio, video, embeddings, or free-form generated text.

API and guardrail integration

Mistral provides direct moderation endpoints for applications that want to submit content and inspect the returned classifications. The model can also back Mistral's custom guardrails. Guardrails are policy controls that can be attached to chat completions, conversations, or agents so that moderation happens as part of a larger request flow.

Guardrail configuration supports category thresholds, selection of the categories to evaluate, blocking actions, and error-handling behavior. When a configured rule is violated, the API can return an HTTP 403 response together with evaluated categories, thresholds, scores, and violation status. This allows an application to centralize enforcement instead of implementing separate moderation checks around every model call.

The model record lists batch API support, but not streaming support or tool and function calling. That matches its role: the output is a structured moderation result rather than a streamed conversational response or an instruction to invoke an external tool. The research identifies structured classification output, while a distinct JSON-mode capability is not verified.

Pricing and availability

Mistral Moderation 2 reached general availability on March 1, 2026. Mistral lists it as an active Premier model and provides access through the Moderations API and batch processing endpoint.

Current Mistral pricing lists input, cached input, and output pricing for this model as free. This makes the model attractive for high-volume screening, where per-request moderation charges could otherwise become a significant part of an application's operating cost. Free listed pricing does not remove other possible requirements such as an API account, service quotas, operational infrastructure, or the need to manage false positives and human review.

Because the endpoint returns classifications and scores instead of ordinary generated prose, Mistral does not publish a conventional maximum output-token limit for this model. The supplied documentation also does not specify a separate output-length limit. The practical output is a moderation result, not a long answer.

Strengths and trade-offs

  • Large moderation context: The 128k-token window is well suited to long documents and multi-turn conversations where earlier context affects the safety judgment.
  • Broad policy coverage: The documented categories span common content-safety risks as well as health, finance, law, PII, and jailbreaking.
  • Direct application integration: Raw-text, conversational, and batch interfaces support both individual checks and higher-volume workflows.
  • Configurable enforcement: Custom guardrails can apply category-specific thresholds and blocking behavior to Mistral chat, conversation, and agent systems.
  • Free listed pricing: Mistral's current pricing page lists input and output pricing as free, which is a meaningful cost advantage for a dedicated moderation layer.

These benefits come with important limitations. Mistral Moderation 2 is not a general-purpose reasoning or generation model, and it is not documented for multimodal moderation. Its category scores require calibration, and a threshold that is too aggressive can block legitimate content while a threshold that is too permissive can miss risky material. Mistral also notes that the moderation endpoint may continue to improve, so applications that depend closely on category scores may need to be recalibrated over time.

Reasoning, coding, and performance characteristics

The model is designed to classify content, not to perform open-ended reasoning or write software. Its practical reasoning capability is therefore limited to the policy classification task represented by its returned categories and scores. It has no documented coding capability, tool use, or function-calling support.

The structured evaluation rates it highly for speed and cost, but those ratings are editorial assessments rather than Mistral-published benchmarks. The cost assessment reflects the provider's current free pricing, while the speed assessment reflects the model's narrow classification role and endpoint design; no independent latency benchmark is supplied here. Compared with a general-purpose language model, a dedicated moderation endpoint is the more appropriate choice when the task is simply to classify safety risks. A general model may be preferable when the system must explain a decision in detail, rewrite content, conduct a complex investigation, or generate a user-facing response.

When to choose Mistral Moderation 2

Choose Mistral Moderation 2 when an application needs a dedicated text-safety layer with broad policy categories, long-context conversation analysis, and configurable enforcement. Typical uses include:

  • screening user-generated posts, messages, or support requests;
  • checking both user prompts and AI-generated replies;
  • identifying possible prompt injection or jailbreaking attempts;
  • moderating long conversations before handing them to an agent;
  • applying category-specific rules to Mistral chat, conversation, or agent workflows;
  • running large-scale moderation checks through the batch endpoint.

It may not be the right choice when the input is primarily an image, recording, or video, because those modalities are not documented for this model. It is also unsuitable as the main model for chat, coding, content generation, embeddings, or complex tool-driven reasoning. In those cases, a general-purpose model or a modality-specific safety system may be more appropriate, with Mistral Moderation 2 used alongside it for text checks where applicable.

Lifecycle and migration considerations

Mistral Moderation 2 replaces Mistral Moderation 2411, which was deprecated on March 31, 2026. New integrations should use mistral-moderation-2603 and the version-two guardrail configuration identified in Mistral's documentation.

Teams migrating from the older model should not assume that existing thresholds will remain optimal. The expanded categories and continuing improvements to the moderation endpoint can change score distributions or enforcement behavior. A sensible migration process is to run representative historical examples through the new model, review false positives and false negatives, then update category thresholds before enabling automatic blocking.

Bottom line

Mistral Moderation 2 is a focused safety-classification service rather than another general-purpose language model. Its strongest reasons to use it are the 128k-token context window, support for text and multi-turn conversations, jailbreaking detection, configurable guardrails, batch processing, and free listed API pricing. Its boundaries are equally clear: it does not generate ordinary text, use tools, or provide documented image, audio, or video moderation. For applications that need a fast, dedicated text policy check, it is a practical fit; for multimodal safety or broader reasoning, it should be paired with other systems rather than used alone.


Answers to Frequently Asked Questions

How much does Mistral Moderation 2 cost and what is its context window?
Mistral's current pricing lists input, cached input, and output pricing for Mistral Moderation 2 as free. The model has a documented 128k-token context window, or 131,072 tokens, allowing it to evaluate long documents and extended conversations.
How can Mistral Moderation 2 be integrated into an application?
Mistral provides Moderations API, conversational, and batch-processing interfaces. The model can also power custom guardrails for chat completions, conversations, and agents, with configurable categories, thresholds, blocking actions, and error handling.
What types of input does Mistral Moderation 2 support?
The model supports text and structured lists of conversational messages, including multi-turn user and assistant exchanges. It is documented as a text moderation model; image, audio, and video input support is not specified.
Can Mistral Moderation 2 detect jailbreak attempts?
Yes. Mistral Moderation 2 includes jailbreaking detection to identify prompts that attempt to manipulate an AI system into bypassing safety instructions, including through role-playing, adversarial wording, or carefully constructed instructions.
What is Mistral Moderation 2 used for?
Mistral Moderation 2 is a text-safety classification model that evaluates raw text and conversational messages for policy risks. Applications can use its category scores and moderation decisions to allow content, trigger human review, display warnings, or block inputs.


Sources 5
Provider

About Mistral AI