What is Mistral Small 4?
Mistral Small 4 is a general-purpose multimodal language model from Mistral AI. The model was released on March 16, 2026, and is listed as active and generally available in the supplied model research. It belongs to the Mistral Small family and is positioned as a hybrid model rather than a system dedicated only to chat, only to reasoning, or only to coding.
In practical terms, Mistral Small 4 can read text and images, reason over information, write and transform text, generate code, and participate in tool-driven workflows. Its output is text: it does not natively generate images, audio, or video. That distinction matters because “multimodal” here describes the model’s ability to accept more than one input type, not its ability to produce every media format.
Mistral AI describes the release as an Apache 2.0 open-weight model. The launch information reports 119 billion total parameters, with approximately 6 billion active parameters per token and 8 billion including embedding and output layers. Current model documentation reports 6.5 billion active parameters, so the exact parameter description depends on which official source and counting convention is used. The important practical point is that the model uses a mixture-of-experts design: only part of the total network is active for each token, which is intended to support a balance between capability and inference efficiency.
Where Mistral Small 4 fits in Mistral AI’s lineup
Mistral Small 4 sits in the smaller-model segment of Mistral AI’s catalog, but its feature set extends beyond basic lightweight text completion. It combines instruction following, reasoning, vision, coding, and agentic capabilities in one model. This makes it a candidate for applications that would otherwise need to route requests among separate chat, vision, coding, and reasoning systems.
The model is especially relevant when deployment cost, response speed, or control over the model matters. Mistral AI provides API access through its inference platform, while the open-weight release supports fine-tuning and self-deployment. Organizations can therefore evaluate it as a hosted model or consider running and adapting it in a more controlled environment, subject to their own infrastructure and licensing review.
Core specifications at a glance
| Specification | Verified detail |
|---|---|
| Provider | Mistral AI |
| Release date | March 16, 2026 |
| Status | Active; generally available |
| Model identifier | mistral-small-2603 |
| Context length | 256,000 tokens |
| Input | Text and images |
| Output | Text |
| Input price | $0.15 per 1 million tokens; cached input $0.015 per 1 million tokens |
| Output price | $0.60 per 1 million tokens |
| Licensing and deployment | Apache 2.0 open-weight release; fine-tuning and self-deployment supported |
The documentation does not provide a verified maximum output-token limit, so applications that require a guaranteed response ceiling should confirm the limit in the specific serving environment before implementation. Mistral AI also does not publish a specific knowledge-cutoff date for this model in the reviewed authoritative sources.
Reasoning and coding capabilities
Mistral Small 4 supports a configurable reasoning_effort setting. This allows an application to trade response depth against speed and resource use. A lower setting may be suitable for routine classification, extraction, or short answers, while a higher setting can be reserved for multi-step analysis, planning, or difficult coding tasks. The supplied research does not establish a universal quality threshold for each setting, so the appropriate configuration should be tested against the workload.
Coding is one of the model’s intended uses. It can generate and explain code, help transform existing code, and support agentic workflows in which a model calls tools or works through a sequence of tasks. The API documentation lists function calling, agents and conversations, built-in tools, structured outputs, predicted outputs, prefix completion, and batching. These features make it more suitable for software assistants and automation pipelines than a model limited to single-turn text generation.
Structured outputs are documented, which can help applications request responses that conform to a defined schema. That is useful for tasks such as extracting invoice fields, classifying support tickets, returning code-review findings, or producing machine-readable workflow steps. A separate legacy JSON-mode capability was not independently verified, so structured outputs should not automatically be described as an equivalent JSON-mode feature.
Images, documents, and tools
Mistral Small 4 accepts image input as well as text. This supports use cases such as examining screenshots, interpreting visual documents, and combining written instructions with image context. The research identifies multimodal document analysis as a primary use case, but it does not provide a detailed list of supported image dimensions, file formats, or image-count limits.
The model also supports document question answering and built-in tools. With the appropriate application layer, it can be used to extract information from long documents, answer questions about supplied material, and contribute to agent workflows that call external functions. Web search support is listed in the broader Mistral ecosystem and in the model research, but web access should be treated as a tool or integration capability rather than as evidence that the base model independently knows current events.
Tool calling does not mean the model performs external actions without safeguards. Production systems still need to validate arguments, control permissions, handle failures, and require approval for sensitive operations. The model can propose or request a tool call; the surrounding application determines whether that call is actually executed.
Pricing, speed, and cost trade-offs
Mistral’s listed inference price is $0.15 per million input tokens and $0.60 per million output tokens. Cached input is priced at $0.015 per million tokens. This makes repeated-context workloads potentially more economical when the serving system can use caching effectively. For example, an application that repeatedly asks questions against a stable instruction prefix or shared reference context may benefit more than a workload consisting entirely of new prompts.
The supplied editorial assessment rates Mistral Small 4 highly for speed and cost, with scores of 8 out of 10 for each, but these are comparative editorial estimates rather than Mistral AI benchmarks. The same assessment gives it a reasoning score of 8, a coding score of 8, and a cost score of 9. These ratings should be read as directional judgments, not provider-published performance guarantees.
The model’s cost advantage may be most meaningful for high-volume workloads where a larger premium model would be unnecessarily expensive. However, lower price does not guarantee lower total cost. Complex tasks may require retries, longer reasoning, additional tool calls, or human review. Teams should measure end-to-end task completion, latency, error rates, and token usage rather than comparing per-token prices alone.
Main strengths and limitations
Key strengths
- Broad capability mix: It combines general instruction following, configurable reasoning, coding, image understanding, and agentic features.
- Large context: The 256,000-token context window is suited to long documents, extensive codebases, and multi-step conversations, within the limits of the particular application.
- Low listed price: Input and output pricing is comparatively economical, with a much lower cached-input rate.
- Open-weight flexibility: Apache 2.0 licensing, fine-tuning, and self-deployment provide options beyond a hosted API.
- Production-oriented interfaces: Function calling, structured outputs, batching, predicted outputs, and document question answering support integration into software systems.
Important limitations
- Text-only output: It does not natively generate images, audio, or video.
- Unknown output ceiling: A documented maximum output-token limit was not verified in the supplied research.
- No published knowledge cutoff: Mistral AI has not provided a specific cutoff date in the reviewed model documentation.
- Reasoning costs more in practice: Deeper reasoning can increase latency and token consumption, even when the per-token price is attractive.
- Probabilistic results: Generated code, extracted data, and tool arguments require validation, especially in workflows that affect external systems.
- Not a speech or media-generation model: Workloads requiring native audio, video, or image creation need a separate specialized system.
Best use cases for Mistral Small 4
Mistral Small 4 is a strong candidate for applications that need several capabilities in one relatively economical model:
- Multimodal document analysis, including questions about text and supplied images.
- Long-context summarization, extraction, and document-based question answering.
- Coding assistants that generate, explain, refactor, or review code.
- Agentic workflows that use function calling, built-in tools, structured responses, or batching.
- High-volume conversational applications where a larger premium model would be unnecessarily costly.
- Private or customized deployments that benefit from open weights, fine-tuning, or greater control over serving infrastructure.
It may be less appropriate when the central requirement is native media generation, specialized speech processing, embeddings, or a guaranteed maximum output length. A larger reasoning-focused model may also be preferable for tasks where the highest available reliability matters more than speed and cost. Conversely, a smaller and simpler model could be a better choice for basic classification or short templated responses that do not need vision, reasoning, or tools.
When to choose Mistral Small 4
Choose Mistral Small 4 when you want one model to cover text and image inputs, coding, configurable reasoning, long-context work, and tool-enabled automation without paying premium-model rates for every request. It is particularly attractive when open-weight licensing, fine-tuning, or self-deployment is part of the requirements.
Choose another option when the task depends on image, audio, or video generation; when a specialized speech or embedding endpoint is required; when a documented output-token maximum is mandatory; or when testing shows that the model’s accuracy on a high-risk task does not justify the savings. The most sensible deployment pattern may also be a routing strategy: use Mistral Small 4 for routine multimodal and coding work, then escalate only difficult or high-consequence requests to a more specialized model.
Bottom line
Mistral Small 4 is a cost-conscious, open-weight model that brings multimodal input, reasoning controls, coding, long context, and tool use into a single general-purpose system. Its strongest distinction is not one isolated benchmark claim, but the combination of broad functionality, low listed inference pricing, and deployment flexibility. Its boundaries are equally clear: it produces text only, lacks a verified published knowledge cutoff and maximum output limit, and still requires application-level validation. For document-heavy, coding, and agentic workloads where capability must be balanced against operating cost, it is a practical model to evaluate.

