What is Mistral Medium 3.5?
Mistral Medium 3.5 is Mistral AI's open-weight model for demanding coding, reasoning, document, and agentic tasks. It is designed to handle workflows in which the model must understand a large amount of context, produce code, call tools, follow instructions over multiple steps, or return structured data that software can process.
The model was released on April 28, 2026, and its canonical API identifier is mistral-medium-3-5. It is available through the Mistral API and as open weights under a Modified MIT license. That combination gives organizations a choice between hosted inference and self-managed deployment, subject to the hardware and operational requirements of a 128-billion-parameter model.
Within Mistral AI's current catalog, Medium 3.5 occupies a high-capability position rather than serving as a lightweight, low-cost model for simple, high-volume generation. Its main focus is the combination of software engineering, long-context analysis, multimodal understanding, reasoning, and tool use.
Technical profile and context window
Mistral Medium 3.5 is a dense model with 128 billion parameters and a 256,000-token context window. A context window is the amount of text, code, documents, tool output, and conversation history that can be supplied to the model for one task. The 256K limit is useful for large repositories, lengthy technical specifications, extended agent traces, and multi-document analysis.
The model accepts text and image inputs. Mistral describes its vision encoder as being trained to handle variable image sizes and aspect ratios, which supports image understanding without requiring every image to use the same dimensions. Supported visual tasks include image interpretation and document question answering. The model's output is text only: it does not natively generate images, audio, or video.
Mistral does not publish a model-specific maximum output-token limit in the supplied documentation. The exact knowledge-cutoff date is also not specified in the current model materials. These are important unknowns for applications that require fixed response budgets or time-sensitive factual coverage.
Coding, reasoning, and agent capabilities
The model is intended for software engineering assistants and coding agents, not only code completion. It can generate and explain code, assist with debugging, reason over project material, and participate in multi-step workflows. Its large context makes it suitable for providing a model with substantial portions of a codebase, documentation, issue history, or agent state at the same time.
Mistral documents configurable reasoning effort through the reasoning_effort parameter. This allows an application to trade deliberation depth against response speed. Lower effort can be appropriate for straightforward transformations or short answers, while higher effort can be reserved for complex analysis, difficult debugging, planning, and longer-running agent tasks. The provider does not publish a single universal reasoning score for the model; any rating of its reasoning capability should therefore be treated as an evaluation judgment rather than a provider specification.
Mistral Medium 3.5 supports function and tool calling. In practical terms, an application can give the model descriptions of available operations and allow it to request those operations in a controlled format. The model is documented for use with Chat Completions, the Agents and Conversations APIs, built-in tools, structured outputs, predicted outputs, prefix support, and batch processing. Tool execution remains an application responsibility unless a higher-level Mistral agent workflow handles it.
Structured outputs are particularly useful when the response must conform to a schema, such as a JSON object containing extracted fields, a task plan, a code-review result, or a classification. This capability should not be confused with a separately verified legacy JSON mode: structured outputs are documented, but an independent model-specific JSON-mode capability was not verified.
Supported inputs and outputs
| Capability | Status | What it means |
|---|---|---|
| Text input | Supported | Accepts prompts, code, documents, conversation history, and tool results. |
| Image input | Supported | Can interpret images and support document question answering. |
| Audio input | Not verified | No audio input support was identified in the supplied model research. |
| Video input | Not verified | No video input support was identified. |
| Text output | Supported | Produces natural-language, code, structured, and tool-related responses. |
| Image, audio, or video output | Not supported | The model is text-output only and does not natively generate these media types. |
Calling the model multimodal refers to its ability to work with more than one input modality, particularly text and images. It does not mean that the model produces every type of media.
Pricing and deployment options
Through the Mistral API, the standard listed price is $1.50 per million input tokens, $0.15 per million cached input tokens, and $7.50 per million output tokens. Cached input pricing can reduce the cost of repeatedly supplying content that the service can reuse, while output generation is substantially more expensive than input processing. Applications should therefore avoid sending unnecessarily long prompts and should control generated response length where possible.
The model is also available as open weights under a Modified MIT license. Mistral's launch material states that self-hosting can be possible with as few as four GPUs, depending on the deployment configuration and quantization. This is a provider claim and should not be interpreted as a universal hardware requirement or guarantee. Actual memory, throughput, latency, and operational costs depend on the selected inference stack, quantization, concurrency, and performance target.
Hosted API access is generally simpler for teams that want managed infrastructure, predictable integration, and no responsibility for serving a large checkpoint. Open-weight deployment may be more appropriate for organizations that need greater control over data handling, network location, customization, or infrastructure, and that can operate the necessary hardware.
Main strengths and trade-offs
- Large working context: The 256,000-token window supports long codebases, extended conversations, large technical documents, and multi-step agent histories.
- Broad task coverage: Coding, reasoning, image understanding, structured responses, and tool use are combined in one model.
- Agent-oriented design: Function calling, built-in tools, Agents and Conversations integration, and configurable reasoning effort support multi-step workflows.
- Deployment control: Open-weight availability can help organizations that need self-hosting or more control over data and infrastructure.
- Cost concentration in output: At $7.50 per million output tokens, extensive generated responses can become expensive compared with smaller models, especially in high-volume workloads.
- Operational demands: A dense 128-billion-parameter checkpoint requires substantial infrastructure for self-hosting, even when quantization is used.
- Text-only generation: It can understand images but cannot replace a dedicated image, audio, or video generation model.
- Unspecified limits: A model-specific maximum output-token limit and knowledge-cutoff date were not found in the supplied authoritative documentation.
The reasoning-effort setting creates a practical speed-versus-depth trade-off. More deliberation may help with difficult coding or planning tasks but can increase latency and token consumption. For short, routine transformations, a smaller or faster model may be more economical.
Best use cases
Mistral Medium 3.5 is a strong fit for coding agents that need to inspect project files, reason about dependencies, draft changes, and return structured plans or patches. It is also suitable for software engineering assistants, code review, technical research, document analysis, and workflows that combine visual material with text instructions.
Examples include analyzing a long software repository alongside its issue history, extracting structured information from image-heavy documents, answering questions across a large collection of technical files, coordinating several tool calls, and producing machine-readable results for a downstream application. Its open-weight availability adds value when deployment control or private infrastructure is more important than minimizing serving complexity.
When to choose Mistral Medium 3.5
Choose Mistral Medium 3.5 when the workload benefits from a large context window, strong coding and reasoning capability, image understanding, tool use, or the option to deploy open weights. It is especially relevant when one model needs to cover several stages of an agentic workflow instead of handing each stage to a separate specialized system.
A smaller model may be more appropriate for simple classification, short rewriting, routine extraction, or very high-volume generation where low cost and low latency matter more than long-context reasoning. A dedicated vision model may be preferable when image processing is the central task, while a media-generation model is required for native image, audio, or video creation. Teams that cannot justify the infrastructure and operational work of serving a large checkpoint may also prefer the hosted API or a smaller managed alternative.
Overall, Mistral Medium 3.5 is best understood as a high-capability, open-weight model for complex text-and-image workflows rather than a universal replacement for every model type. Its main value comes from bringing long context, coding, reasoning, structured responses, and tool interaction together, while its principal costs are API output pricing, hardware demands, and the absence of native media generation.

