Ministral 3

Ministral 3 8B

by Mistral AI · Active; generally available

Ministral 3 8B is an Apache 2.0 open-weight model designed for efficient local, edge, and hosted inference. It accepts text and image inputs, produces text, supports function calling, structured outputs, document question answering, batching, and a 256k-token context window. Hosted pricing is $0.15 per million input tokens and $0.15 per million output tokens, with cached input priced at $0.015 per million tokens.

Text Reasoning Coding
Ministral 3 8B is Mistral AI's compact multimodal model for developers who need text and image understanding without the infrastructure demands or cost of a much larger model. Released on December 2, 2025, it is aimed at local and edge deployment as well as low-cost hosted API workloads. The model supports document question answering, function calling, structured outputs, batching, and a 256k-token context window, while its Apache 2.0 license permits broad deployment options subject to the license terms.
Outputs

What Ministral 3 8B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Ministral 3
Model type Lightweight
Context window 256K tokens
Release date 2025-12-02
Status Active; generally available
Knowledge cutoff notes

No explicit knowledge-cutoff date was identified in the reviewed first-party model documentation or lifecycle material.

Model notes

The canonical dated API model ID is ministral-8b-2512, while the rolling hosted identifier is ministral-8b-latest. Mistral lists base, instruct, and reasoning variants within the Ministral 3 8B family. The model is Apache 2.0, open-weight, and designed for local or edge deployment. Current first-party documentation lists text and vision capabilities, structured outputs, function calling, document QnA, prefix completion, chat completions, and batching. The exact maximum output-token limit, knowledge cutoff, fine-tuning availability, and separate legacy JSON-mode support were not directly verified in the reviewed first-party sources. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input $0.15 per million tokens; cached input $0.015 per million tokens
Output $0.15 per million tokens
Model guide

Ministral 3 8B: Efficient Vision-Language Inference at the Edge

Ministral 3 8B is an Apache 2.0 open-weight model from Mistral AI designed for efficient local, edge, and hosted inference. It accepts text and image inputs, produces text, supports function calling and structured outputs, offers a 256k-token context window, and is priced at $0.15 per million input and output tokens.

What Ministral 3 8B is

Ministral 3 8B is an 8-billion-parameter-class, open-weight language model from Mistral AI. It belongs to the Ministral 3 family, which is positioned around efficient inference on diverse hardware rather than maximum capability at any cost. The model is available for hosted inference through Mistral and is also intended for local or edge deployment using the weights and software stack selected by the developer.

The model is released under the Apache 2.0 license. That makes it a notably flexible option for organizations that need to run inference in their own environment, keep data close to its source, or avoid depending entirely on a hosted endpoint. Local deployment still requires suitable hardware, memory, quantization, and runtime support, so the license does not by itself guarantee that every device can run the model comfortably.

The canonical dated API model ID is ministral-8b-2512; Mistral also provides the rolling hosted identifier ministral-8b-latest. Mistral describes base, instruct, and reasoning variants within the broader Ministral 3 8B family, so users should confirm the exact identifier and variant before deploying an application.

Modalities and core capabilities

Ministral 3 8B accepts text and image inputs and produces text output. Its image capability is for understanding visual content: an application can submit an image alongside a question or instruction and ask the model to describe, extract, compare, or reason about information visible in that image. It is not an image-generation model, and the supplied specifications do not identify native audio or video input or output.

The model supports document question answering, which makes it useful for extracting information from documents and answering questions grounded in supplied content. Image understanding can extend this to scanned pages, diagrams, screenshots, or other visual material, although the quality of results will depend on image clarity and the complexity of the content.

Function calling lets the model request an application-defined operation in a structured form. For example, a developer could expose a document lookup, database query, or workflow action and allow Ministral 3 8B to select and populate that function. The model does not execute external actions by itself; the surrounding application must validate the request and perform the operation.

Structured outputs are also supported. This is useful when the response needs to follow a defined schema, such as a list of extracted fields, a classification result, or a document-processing record. Structured outputs should not automatically be treated as a separate legacy JSON-mode feature. The exact parameters and endpoint behavior should be checked against the API or runtime being used.

Context window and output limits

Mistral's current documentation lists a 256,000-token context window for Ministral 3 8B. A context window is the total amount of text and other tokenized input that can be considered in one request, along with the generated response and any system or tool messages counted by the serving implementation. In practical terms, the large window can accommodate long documents, extended conversations, and sizeable retrieval results.

A large context window does not mean that every deployment will process very long requests quickly or cheaply. Memory requirements, quantization, hardware, concurrent requests, and the API's own request limits affect real-world performance. Long prompts also increase input-token charges when using hosted inference.

The reviewed first-party material does not specify a maximum output-token limit. Developers should therefore treat the output ceiling as deployment-specific until it is confirmed in the documentation for the selected endpoint, model variant, or local runtime. The absence of a verified limit is different from saying that the model has no output limit.

Pricing and deployment options

Mistral's current hosted pricing lists Ministral 3 8B at $0.15 per million input tokens and $0.15 per million output tokens. Cached input is listed at $0.015 per million tokens. These are usage prices rather than a subscription fee, and the input and output totals should be calculated separately for each workload.

At that rate, the model is suited to high-volume extraction, classification, routing, and generation tasks where a larger model would add cost without providing enough additional value. A workload that sends large documents repeatedly should account for input volume even when the generated answers are short. Cached-input pricing may reduce the cost of repeated prompt material where the applicable serving mechanism supports it.

Hosted inference is only one deployment path. Because Ministral 3 8B is open-weight and designed for local and edge use, an organization can instead manage inference in its own environment. That can improve data locality and give the operator more control over latency and availability, but it transfers responsibility for hardware, scaling, monitoring, upgrades, and model-runtime compatibility to the operator.

Speed, cost, and capability trade-offs

The main reason to choose Ministral 3 8B is efficiency. An 8B-class model is generally easier to host and less expensive to serve than substantially larger models, and Mistral positions this family for diverse hardware and edge scenarios. The supplied comparative editorial assessment rates its speed and cost highly, with a speed score of 9 and cost score of 9. Those scores are editorial estimates, not vendor-published benchmark results.

The trade-off is depth. The same assessment gives the model a reasoning score of 7 and coding score of 7, which indicates a useful middle position rather than frontier-level performance. These scores should not be interpreted as standardized benchmark results. In difficult multi-step reasoning, advanced software engineering, or highly specialized analysis, a larger model may produce more reliable results at the expense of speed, hardware requirements, or price.

Ministral 3 8B is better understood as a practical general-purpose model for everyday generation, multimodal extraction, tool routing, and edge inference. It can support coding-related workflows through text generation and function calling, but the supplied research does not verify a specialized coding mode or a particular coding benchmark result. Code produced by the model should be tested and reviewed like code from any other probabilistic system.

Reasoning, tools, and structured workflows

The wider Ministral 3 8B family includes base, instruct, and reasoning variants. The reasoning label describes a model-family distinction, so users should verify which variant they are calling rather than assume that every family identifier has identical behavior. The reviewed documentation does not provide a standardized reasoning benchmark or a guaranteed performance level for complex problems.

For tool-using applications, function calling and structured outputs are more important than a marketing category. A lightweight agent can ask the model to classify an incoming request, select a permitted tool, produce arguments in a defined schema, and then use the tool result to continue the workflow. Validation, permissions, error handling, and confirmation steps should remain in application code, especially when tools can change records or trigger external actions.

Batch processing is supported, which can be useful when many independent documents, images, or classification requests can be processed without interactive latency requirements. Streaming is also listed as supported in the supplied model data, allowing applications to display generated text progressively when the selected endpoint exposes that behavior.

Best use cases for Ministral 3 8B

  • Edge and local assistants: Applications that need text interaction and image understanding close to the user or device.
  • Document workflows: Question answering, field extraction, classification, and structured records from text-heavy or image-based documents.
  • Visual question answering: Analyzing screenshots, scans, diagrams, and other supplied images with natural-language instructions.
  • Lightweight agents: Tool selection, routing, structured action requests, and workflow coordination where a compact model is sufficient.
  • Private or offline processing: Deployments where sending data to a hosted service is undesirable and local infrastructure is available.
  • High-volume generation: Summarization, transformation, classification, and other token-intensive tasks where hosted cost and latency matter.
  • Long-context retrieval workflows: Applications that need to provide substantial document collections or conversation history in one request.

When to choose this model

Choose Ministral 3 8B when deployment flexibility, low token cost, and responsive inference matter more than maximizing performance on the hardest reasoning tasks. It is particularly attractive when the application needs image understanding, structured responses, and tool calling in the same compact model, or when an organization wants the option of local, private, or edge deployment.

It is also a sensible starting point for a document pipeline. A system could submit a document image or text, ask targeted questions, and request the result in a defined schema. The 256k-token context window provides room for long source material, although developers should still test retrieval and chunking strategies rather than assume that placing everything in one prompt is always optimal.

A larger model may be more appropriate for difficult mathematical reasoning, complex planning, advanced coding, or tasks where occasional errors have a high operational cost. A hosted-only model may be preferable when the team does not want to manage local inference infrastructure. A specialized vision, audio, or generative-image system is required when the application needs those output modalities, because Ministral 3 8B produces text and does not natively generate images, audio, or video.

Limitations and unverified details

Ministral 3 8B's compact size is an efficiency advantage, but it also limits how much capability should be expected from it. The model may be less dependable than larger systems on long chains of reasoning, highly advanced coding, and specialized professional tasks. Outputs remain probabilistic and should be checked, especially when they influence decisions or invoke tools.

The model does not provide a verified built-in source of current information. Applications that require fresh facts should supply retrieval, search, or another independently verified data source. The reviewed first-party material does not state a knowledge-cutoff date, so no specific cutoff should be assumed.

The maximum output-token limit, fine-tuning availability, and separate legacy JSON-mode support were not directly verified in the reviewed sources. These details may vary by endpoint or deployment. Confirm them before committing to a production design.

Bottom line

Ministral 3 8B is a practical open-weight model for developers who need efficient text generation and image understanding rather than the highest possible reasoning ceiling. Its combination of Apache 2.0 licensing, local and hosted deployment options, 256k-token context, function calling, structured outputs, batching, and low hosted pricing makes it well suited to edge assistants, document processing, visual extraction, and lightweight agent workflows. Its main limitations are the expected capability gap versus larger models, the lack of native media generation, and several deployment details that must be verified for the specific API or runtime.


Answers to Frequently Asked Questions

What are the best use cases for Ministral 3 8B?
Suitable use cases include edge assistants, private or offline processing, document question answering, visual question answering, structured data extraction, high-volume classification and generation, long-context retrieval workflows, and lightweight agents using function calling. Larger models may be preferable for advanced reasoning, complex planning, or highly specialized coding.
Can Ministral 3 8B run locally or at the edge?
Yes. Ministral 3 8B is open-weight, licensed under Apache 2.0, and intended for local and edge deployment as well as hosted inference. Local operation still depends on suitable hardware, memory, quantization, and runtime compatibility.
What is the context window and pricing for Ministral 3 8B?
Mistral lists a 256,000-token context window for Ministral 3 8B. Hosted pricing is $0.15 per million input tokens and $0.15 per million output tokens, with cached input listed at $0.015 per million tokens. Actual performance and limits can vary by endpoint or local runtime.
What is Ministral 3 8B?
Ministral 3 8B is an 8-billion-parameter open-weight language model from Mistral AI designed for efficient text generation, image understanding, document processing, and edge or local deployment. It is available under the Apache 2.0 license and can also be accessed through hosted inference.
What modalities does Ministral 3 8B support?
Ministral 3 8B accepts text and image inputs and produces text outputs. It can analyze documents, screenshots, scans, diagrams, and other images, but it is not an image-generation model, and native audio or video support is not specified.


Sources 6
Provider

About Mistral AI