Mistral Small

Mistral Small 4

by Mistral AI · Active; generally available

Mistral Small 4 is an Apache 2.0 open-weight model for text and image understanding, reasoning, coding, document analysis, and agentic workflows. It offers a 256,000-token context window, configurable reasoning effort, function calling, structured outputs, batching, fine-tuning, and self-deployment. API pricing is $0.15 per million input tokens, $0.015 for cached input, and $0.60 per million output tokens. The model produces text only and has no verified published maximum output-token limit or knowledge-cutoff date.

Text Reasoning Coding
Mistral Small 4 is designed for applications that need more than straightforward text generation but cannot justify the cost or latency of a larger premium model. It accepts text and images, supports configurable reasoning, coding, function calling, built-in tools, structured outputs, batching, and fine-tuning, and can be self-deployed from its open-weight release. At $0.15 per million input tokens and $0.60 per million output tokens, it targets cost-sensitive production workloads such as document analysis, coding assistants, multimodal chat, and agentic automation.
Outputs

What Mistral Small 4 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Fine-tuning Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Mistral Small
Model type Multimodal
Context window 256K tokens
Release date 2026-03-16
Status Active; generally available
Knowledge cutoff notes

Mistral AI does not publish a specific knowledge-cutoff date for Mistral Small 4 in the authoritative model documentation or launch announcement reviewed.

Model notes

Mistral Small 4 is an Apache 2.0 open-weight hybrid mixture-of-experts model. The official announcement describes 119B total parameters, approximately 6B active parameters per token, and 8B including embedding and output layers; the current model documentation reports 6.5B active parameters. It combines instruct, reasoning, multimodal, coding, and agentic capabilities and supports a configurable reasoning_effort setting. The canonical API identifier is mistral-small-2603. Official documentation lists chat completions, function calling, agents and conversations, built-in tools, structured outputs, predicted outputs, document QnA, prefix completion, and batching. Fine-tuning and self-deployment are supported through the open-weight release. JSON mode as a separate legacy capability was not independently verified; structured outputs are documented. The model accepts text and image inputs and produces text responses. Editorial scores are comparative estimates, not vendor-provided ratings.

Cost

Model pricing

Input $0.15 per 1M tokens; cached input $0.015 per 1M tokens
Output $0.60 per 1M tokens
Model guide

Mistral Small 4: An Open-Weight Multimodal Model for Fast, Low-Cost Reasoning

Mistral Small 4 is Mistral AI’s open-weight hybrid mixture-of-experts model for general chat, document and image understanding, coding, reasoning, and agentic workflows. It combines a 256,000-token context window with relatively low API pricing, configurable reasoning, tool support, and Apache 2.0 licensing, while producing text rather than images, audio, or video.

What is Mistral Small 4?

Mistral Small 4 is a general-purpose multimodal language model from Mistral AI. The model was released on March 16, 2026, and is listed as active and generally available in the supplied model research. It belongs to the Mistral Small family and is positioned as a hybrid model rather than a system dedicated only to chat, only to reasoning, or only to coding.

In practical terms, Mistral Small 4 can read text and images, reason over information, write and transform text, generate code, and participate in tool-driven workflows. Its output is text: it does not natively generate images, audio, or video. That distinction matters because “multimodal” here describes the model’s ability to accept more than one input type, not its ability to produce every media format.

Mistral AI describes the release as an Apache 2.0 open-weight model. The launch information reports 119 billion total parameters, with approximately 6 billion active parameters per token and 8 billion including embedding and output layers. Current model documentation reports 6.5 billion active parameters, so the exact parameter description depends on which official source and counting convention is used. The important practical point is that the model uses a mixture-of-experts design: only part of the total network is active for each token, which is intended to support a balance between capability and inference efficiency.

Where Mistral Small 4 fits in Mistral AI’s lineup

Mistral Small 4 sits in the smaller-model segment of Mistral AI’s catalog, but its feature set extends beyond basic lightweight text completion. It combines instruction following, reasoning, vision, coding, and agentic capabilities in one model. This makes it a candidate for applications that would otherwise need to route requests among separate chat, vision, coding, and reasoning systems.

The model is especially relevant when deployment cost, response speed, or control over the model matters. Mistral AI provides API access through its inference platform, while the open-weight release supports fine-tuning and self-deployment. Organizations can therefore evaluate it as a hosted model or consider running and adapting it in a more controlled environment, subject to their own infrastructure and licensing review.

Core specifications at a glance

SpecificationVerified detail
ProviderMistral AI
Release dateMarch 16, 2026
StatusActive; generally available
Model identifiermistral-small-2603
Context length256,000 tokens
InputText and images
OutputText
Input price$0.15 per 1 million tokens; cached input $0.015 per 1 million tokens
Output price$0.60 per 1 million tokens
Licensing and deploymentApache 2.0 open-weight release; fine-tuning and self-deployment supported

The documentation does not provide a verified maximum output-token limit, so applications that require a guaranteed response ceiling should confirm the limit in the specific serving environment before implementation. Mistral AI also does not publish a specific knowledge-cutoff date for this model in the reviewed authoritative sources.

Reasoning and coding capabilities

Mistral Small 4 supports a configurable reasoning_effort setting. This allows an application to trade response depth against speed and resource use. A lower setting may be suitable for routine classification, extraction, or short answers, while a higher setting can be reserved for multi-step analysis, planning, or difficult coding tasks. The supplied research does not establish a universal quality threshold for each setting, so the appropriate configuration should be tested against the workload.

Coding is one of the model’s intended uses. It can generate and explain code, help transform existing code, and support agentic workflows in which a model calls tools or works through a sequence of tasks. The API documentation lists function calling, agents and conversations, built-in tools, structured outputs, predicted outputs, prefix completion, and batching. These features make it more suitable for software assistants and automation pipelines than a model limited to single-turn text generation.

Structured outputs are documented, which can help applications request responses that conform to a defined schema. That is useful for tasks such as extracting invoice fields, classifying support tickets, returning code-review findings, or producing machine-readable workflow steps. A separate legacy JSON-mode capability was not independently verified, so structured outputs should not automatically be described as an equivalent JSON-mode feature.

Images, documents, and tools

Mistral Small 4 accepts image input as well as text. This supports use cases such as examining screenshots, interpreting visual documents, and combining written instructions with image context. The research identifies multimodal document analysis as a primary use case, but it does not provide a detailed list of supported image dimensions, file formats, or image-count limits.

The model also supports document question answering and built-in tools. With the appropriate application layer, it can be used to extract information from long documents, answer questions about supplied material, and contribute to agent workflows that call external functions. Web search support is listed in the broader Mistral ecosystem and in the model research, but web access should be treated as a tool or integration capability rather than as evidence that the base model independently knows current events.

Tool calling does not mean the model performs external actions without safeguards. Production systems still need to validate arguments, control permissions, handle failures, and require approval for sensitive operations. The model can propose or request a tool call; the surrounding application determines whether that call is actually executed.

Pricing, speed, and cost trade-offs

Mistral’s listed inference price is $0.15 per million input tokens and $0.60 per million output tokens. Cached input is priced at $0.015 per million tokens. This makes repeated-context workloads potentially more economical when the serving system can use caching effectively. For example, an application that repeatedly asks questions against a stable instruction prefix or shared reference context may benefit more than a workload consisting entirely of new prompts.

The supplied editorial assessment rates Mistral Small 4 highly for speed and cost, with scores of 8 out of 10 for each, but these are comparative editorial estimates rather than Mistral AI benchmarks. The same assessment gives it a reasoning score of 8, a coding score of 8, and a cost score of 9. These ratings should be read as directional judgments, not provider-published performance guarantees.

The model’s cost advantage may be most meaningful for high-volume workloads where a larger premium model would be unnecessarily expensive. However, lower price does not guarantee lower total cost. Complex tasks may require retries, longer reasoning, additional tool calls, or human review. Teams should measure end-to-end task completion, latency, error rates, and token usage rather than comparing per-token prices alone.

Main strengths and limitations

Key strengths

  • Broad capability mix: It combines general instruction following, configurable reasoning, coding, image understanding, and agentic features.
  • Large context: The 256,000-token context window is suited to long documents, extensive codebases, and multi-step conversations, within the limits of the particular application.
  • Low listed price: Input and output pricing is comparatively economical, with a much lower cached-input rate.
  • Open-weight flexibility: Apache 2.0 licensing, fine-tuning, and self-deployment provide options beyond a hosted API.
  • Production-oriented interfaces: Function calling, structured outputs, batching, predicted outputs, and document question answering support integration into software systems.

Important limitations

  • Text-only output: It does not natively generate images, audio, or video.
  • Unknown output ceiling: A documented maximum output-token limit was not verified in the supplied research.
  • No published knowledge cutoff: Mistral AI has not provided a specific cutoff date in the reviewed model documentation.
  • Reasoning costs more in practice: Deeper reasoning can increase latency and token consumption, even when the per-token price is attractive.
  • Probabilistic results: Generated code, extracted data, and tool arguments require validation, especially in workflows that affect external systems.
  • Not a speech or media-generation model: Workloads requiring native audio, video, or image creation need a separate specialized system.

Best use cases for Mistral Small 4

Mistral Small 4 is a strong candidate for applications that need several capabilities in one relatively economical model:

  • Multimodal document analysis, including questions about text and supplied images.
  • Long-context summarization, extraction, and document-based question answering.
  • Coding assistants that generate, explain, refactor, or review code.
  • Agentic workflows that use function calling, built-in tools, structured responses, or batching.
  • High-volume conversational applications where a larger premium model would be unnecessarily costly.
  • Private or customized deployments that benefit from open weights, fine-tuning, or greater control over serving infrastructure.

It may be less appropriate when the central requirement is native media generation, specialized speech processing, embeddings, or a guaranteed maximum output length. A larger reasoning-focused model may also be preferable for tasks where the highest available reliability matters more than speed and cost. Conversely, a smaller and simpler model could be a better choice for basic classification or short templated responses that do not need vision, reasoning, or tools.

When to choose Mistral Small 4

Choose Mistral Small 4 when you want one model to cover text and image inputs, coding, configurable reasoning, long-context work, and tool-enabled automation without paying premium-model rates for every request. It is particularly attractive when open-weight licensing, fine-tuning, or self-deployment is part of the requirements.

Choose another option when the task depends on image, audio, or video generation; when a specialized speech or embedding endpoint is required; when a documented output-token maximum is mandatory; or when testing shows that the model’s accuracy on a high-risk task does not justify the savings. The most sensible deployment pattern may also be a routing strategy: use Mistral Small 4 for routine multimodal and coding work, then escalate only difficult or high-consequence requests to a more specialized model.

Bottom line

Mistral Small 4 is a cost-conscious, open-weight model that brings multimodal input, reasoning controls, coding, long context, and tool use into a single general-purpose system. Its strongest distinction is not one isolated benchmark claim, but the combination of broad functionality, low listed inference pricing, and deployment flexibility. Its boundaries are equally clear: it produces text only, lacks a verified published knowledge cutoff and maximum output limit, and still requires application-level validation. For document-heavy, coding, and agentic workloads where capability must be balanced against operating cost, it is a practical model to evaluate.


Answers to Frequently Asked Questions

What are the limitations of Mistral Small 4?
Mistral Small 4 produces text only and does not natively generate images, audio, or video. Its documentation does not verify a maximum output-token limit or a specific knowledge cutoff, and generated code, extracted data, and tool arguments should be validated before use in production.
Is Mistral Small 4 open source, and can it be self-hosted?
Mistral Small 4 is released as an Apache 2.0 open-weight model. The release supports fine-tuning and self-deployment, while Mistral AI also provides hosted API access through its inference platform.
What are the main capabilities of Mistral Small 4?
Mistral Small 4 supports text and image understanding, a 256,000-token context window, configurable reasoning effort, code generation and transformation, document question answering, structured outputs, function calling, batching, and tool-enabled workflows.
What is Mistral Small 4?
Mistral Small 4 is a general-purpose multimodal language model from Mistral AI that accepts text and images and produces text. It supports configurable reasoning, coding, long-context processing, structured outputs, function calling, and agentic workflows.
How much does Mistral Small 4 cost?
The listed price is $0.15 per 1 million input tokens, $0.015 per 1 million cached input tokens, and $0.60 per 1 million output tokens. Actual total cost can also depend on reasoning depth, retries, tool calls, and human review.


Sources 7
Provider

About Mistral AI