Mistral Medium

Mistral Medium 3.5

by Mistral AI · GA; currently available; open weights

Mistral Medium 3.5 is a 128-billion-parameter open-weight model for long-context coding, reasoning, image understanding, tool use and structured outputs. It accepts text and images, supports a 256,000-token context window and configurable reasoning effort, and costs $1.50 per million input tokens plus $7.50 per million output tokens through the Mistral API.

Text Reasoning Coding
Mistral Medium 3.5 is a frontier-class multimodal model designed for long-horizon agentic workflows, software engineering, tool calling, and structured machine-readable responses. It is available through the Mistral API and as open weights under a Modified MIT license.
Outputs

What Mistral Medium 3.5 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
7/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Mistral Medium
Model type Multimodal
Context window 256K tokens
Release date 2026-04-28
Status GA; currently available; open weights
Knowledge cutoff notes

No exact model-specific knowledge-cutoff date was identified in the current Mistral model documentation or launch material.

Model notes

Canonical API identifier is mistral-medium-3-5. The model is a dense 128B checkpoint with a 256K-token context window and a vision encoder for variable-size image inputs. Mistral documents configurable reasoning effort through the reasoning_effort parameter. It is available through Chat Completions, Agents and Conversations, built-in tools, function calling, structured outputs, predicted outputs, prefix support, document question answering and batching. The model is released as open weights under a Modified MIT license. Mistral's launch material reports that self-hosting is possible with as few as four GPUs, depending on deployment and quantization. Reasoning, coding, speed and cost values are editorial comparative estimates rather than provider-published scores. A model-specific knowledge-cutoff date and maximum output-token limit were not found in the current authoritative documentation. Structured outputs are documented, but a separate legacy JSON-mode capability was not independently verified.

Cost

Model pricing

Input $1.50 per million input tokens; $0.15 per million cached input tokens
Output $7.50 per million output tokens
Model guide

Mistral Medium 3.5: Open-Weight Model for Long-Context Agentic Coding

Mistral Medium 3.5 is a 128-billion-parameter open-weight multimodal model released by Mistral AI on April 28, 2026. It combines instruction following, reasoning, coding, vision, structured outputs, and tool use in a single model with a 256,000-token context window.

What is Mistral Medium 3.5?

Mistral Medium 3.5 is Mistral AI's open-weight model for demanding coding, reasoning, document, and agentic tasks. It is designed to handle workflows in which the model must understand a large amount of context, produce code, call tools, follow instructions over multiple steps, or return structured data that software can process.

The model was released on April 28, 2026, and its canonical API identifier is mistral-medium-3-5. It is available through the Mistral API and as open weights under a Modified MIT license. That combination gives organizations a choice between hosted inference and self-managed deployment, subject to the hardware and operational requirements of a 128-billion-parameter model.

Within Mistral AI's current catalog, Medium 3.5 occupies a high-capability position rather than serving as a lightweight, low-cost model for simple, high-volume generation. Its main focus is the combination of software engineering, long-context analysis, multimodal understanding, reasoning, and tool use.

Technical profile and context window

Mistral Medium 3.5 is a dense model with 128 billion parameters and a 256,000-token context window. A context window is the amount of text, code, documents, tool output, and conversation history that can be supplied to the model for one task. The 256K limit is useful for large repositories, lengthy technical specifications, extended agent traces, and multi-document analysis.

The model accepts text and image inputs. Mistral describes its vision encoder as being trained to handle variable image sizes and aspect ratios, which supports image understanding without requiring every image to use the same dimensions. Supported visual tasks include image interpretation and document question answering. The model's output is text only: it does not natively generate images, audio, or video.

Mistral does not publish a model-specific maximum output-token limit in the supplied documentation. The exact knowledge-cutoff date is also not specified in the current model materials. These are important unknowns for applications that require fixed response budgets or time-sensitive factual coverage.

Coding, reasoning, and agent capabilities

The model is intended for software engineering assistants and coding agents, not only code completion. It can generate and explain code, assist with debugging, reason over project material, and participate in multi-step workflows. Its large context makes it suitable for providing a model with substantial portions of a codebase, documentation, issue history, or agent state at the same time.

Mistral documents configurable reasoning effort through the reasoning_effort parameter. This allows an application to trade deliberation depth against response speed. Lower effort can be appropriate for straightforward transformations or short answers, while higher effort can be reserved for complex analysis, difficult debugging, planning, and longer-running agent tasks. The provider does not publish a single universal reasoning score for the model; any rating of its reasoning capability should therefore be treated as an evaluation judgment rather than a provider specification.

Mistral Medium 3.5 supports function and tool calling. In practical terms, an application can give the model descriptions of available operations and allow it to request those operations in a controlled format. The model is documented for use with Chat Completions, the Agents and Conversations APIs, built-in tools, structured outputs, predicted outputs, prefix support, and batch processing. Tool execution remains an application responsibility unless a higher-level Mistral agent workflow handles it.

Structured outputs are particularly useful when the response must conform to a schema, such as a JSON object containing extracted fields, a task plan, a code-review result, or a classification. This capability should not be confused with a separately verified legacy JSON mode: structured outputs are documented, but an independent model-specific JSON-mode capability was not verified.

Supported inputs and outputs

CapabilityStatusWhat it means
Text inputSupportedAccepts prompts, code, documents, conversation history, and tool results.
Image inputSupportedCan interpret images and support document question answering.
Audio inputNot verifiedNo audio input support was identified in the supplied model research.
Video inputNot verifiedNo video input support was identified.
Text outputSupportedProduces natural-language, code, structured, and tool-related responses.
Image, audio, or video outputNot supportedThe model is text-output only and does not natively generate these media types.

Calling the model multimodal refers to its ability to work with more than one input modality, particularly text and images. It does not mean that the model produces every type of media.

Pricing and deployment options

Through the Mistral API, the standard listed price is $1.50 per million input tokens, $0.15 per million cached input tokens, and $7.50 per million output tokens. Cached input pricing can reduce the cost of repeatedly supplying content that the service can reuse, while output generation is substantially more expensive than input processing. Applications should therefore avoid sending unnecessarily long prompts and should control generated response length where possible.

The model is also available as open weights under a Modified MIT license. Mistral's launch material states that self-hosting can be possible with as few as four GPUs, depending on the deployment configuration and quantization. This is a provider claim and should not be interpreted as a universal hardware requirement or guarantee. Actual memory, throughput, latency, and operational costs depend on the selected inference stack, quantization, concurrency, and performance target.

Hosted API access is generally simpler for teams that want managed infrastructure, predictable integration, and no responsibility for serving a large checkpoint. Open-weight deployment may be more appropriate for organizations that need greater control over data handling, network location, customization, or infrastructure, and that can operate the necessary hardware.

Main strengths and trade-offs

  • Large working context: The 256,000-token window supports long codebases, extended conversations, large technical documents, and multi-step agent histories.
  • Broad task coverage: Coding, reasoning, image understanding, structured responses, and tool use are combined in one model.
  • Agent-oriented design: Function calling, built-in tools, Agents and Conversations integration, and configurable reasoning effort support multi-step workflows.
  • Deployment control: Open-weight availability can help organizations that need self-hosting or more control over data and infrastructure.
  • Cost concentration in output: At $7.50 per million output tokens, extensive generated responses can become expensive compared with smaller models, especially in high-volume workloads.
  • Operational demands: A dense 128-billion-parameter checkpoint requires substantial infrastructure for self-hosting, even when quantization is used.
  • Text-only generation: It can understand images but cannot replace a dedicated image, audio, or video generation model.
  • Unspecified limits: A model-specific maximum output-token limit and knowledge-cutoff date were not found in the supplied authoritative documentation.

The reasoning-effort setting creates a practical speed-versus-depth trade-off. More deliberation may help with difficult coding or planning tasks but can increase latency and token consumption. For short, routine transformations, a smaller or faster model may be more economical.

Best use cases

Mistral Medium 3.5 is a strong fit for coding agents that need to inspect project files, reason about dependencies, draft changes, and return structured plans or patches. It is also suitable for software engineering assistants, code review, technical research, document analysis, and workflows that combine visual material with text instructions.

Examples include analyzing a long software repository alongside its issue history, extracting structured information from image-heavy documents, answering questions across a large collection of technical files, coordinating several tool calls, and producing machine-readable results for a downstream application. Its open-weight availability adds value when deployment control or private infrastructure is more important than minimizing serving complexity.

When to choose Mistral Medium 3.5

Choose Mistral Medium 3.5 when the workload benefits from a large context window, strong coding and reasoning capability, image understanding, tool use, or the option to deploy open weights. It is especially relevant when one model needs to cover several stages of an agentic workflow instead of handing each stage to a separate specialized system.

A smaller model may be more appropriate for simple classification, short rewriting, routine extraction, or very high-volume generation where low cost and low latency matter more than long-context reasoning. A dedicated vision model may be preferable when image processing is the central task, while a media-generation model is required for native image, audio, or video creation. Teams that cannot justify the infrastructure and operational work of serving a large checkpoint may also prefer the hosted API or a smaller managed alternative.

Overall, Mistral Medium 3.5 is best understood as a high-capability, open-weight model for complex text-and-image workflows rather than a universal replacement for every model type. Its main value comes from bringing long context, coding, reasoning, structured responses, and tool interaction together, while its principal costs are API output pricing, hardware demands, and the absence of native media generation.


Answers to Frequently Asked Questions

Can Mistral Medium 3.5 be self-hosted?
Yes. Mistral Medium 3.5 is available as open weights under a Modified MIT license. Mistral states that self-hosting may be possible with as few as four GPUs depending on the deployment configuration and quantization, but actual hardware, memory, latency, and operating costs vary.
How much does Mistral Medium 3.5 cost through the API?
The listed API price is $1.50 per million input tokens, $0.15 per million cached input tokens, and $7.50 per million output tokens. Because output tokens are considerably more expensive, applications should control response length and avoid sending unnecessary context.
How large is Mistral Medium 3.5's context window?
Mistral Medium 3.5 has a 256,000-token context window. This allows it to process large codebases, lengthy technical documents, extended conversations, tool outputs, and multi-step agent histories in a single workflow.
What can Mistral Medium 3.5 be used for?
The model is suited to coding agents, software engineering assistants, debugging, code review, technical research, document question answering, image-and-text analysis, structured data extraction, and multi-step workflows involving tool calls.
What is Mistral Medium 3.5?
Mistral Medium 3.5 is a 128-billion-parameter open-weight model designed for coding, reasoning, long-context document analysis, image understanding, structured outputs, and agentic workflows. It is available through the Mistral API and under a Modified MIT license for self-managed deployment.


Sources 6
Provider

About Mistral AI