Ministral 3

Ministral 3 14B

by Mistral AI · Active; generally available

Ministral 3 14B is Mistral AI’s largest compact Ministral model, offering text and image understanding, function calling, structured outputs, a 256k-token context window, and open-weight local deployment under the Apache 2.0 license.

Text Reasoning Coding
Ministral 3 14B is a 14-billion-parameter open-weight multimodal model released by Mistral AI on December 2, 2025. It is the largest member of the Ministral 3 family and targets developers who need advanced language and vision capabilities while retaining the option to deploy the model on local hardware.
Outputs

What Ministral 3 14B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Fine-tuning Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Ministral 3
Model type Multimodal
Context window 262K tokens
Release date 2025-12-02
Status Active; generally available
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was identified in the current Mistral model documentation or model card.

Model notes

The canonical hosted API identifier is ministral-14b-2512. Mistral tooling also exposes ministral-14b-latest as a rolling pointer, which is not a separate model entity. The Ministral 3 14B family includes Base, Instruct, and Reasoning checkpoints; the hosted model page describes the general model while the open-weight instruct checkpoint is specifically optimized for instruction-following. The instruct model combines a 13.5B language model with a 0.4B vision encoder and is available in FP8, with BF16 and other community or official checkpoint formats also available. Mistral documents approximately 24 GB of VRAM for the FP8 instruct checkpoint, with lower requirements possible through quantization. Fine-tuning is supported through the availability of the open-weight base checkpoint, but this should not be interpreted as unrestricted hosted API fine-tuning for every deployment.

Cost

Model pricing

Input $0.20 per 1 million tokens; cached input $0.02 per 1 million tokens
Output $0.20 per 1 million tokens
Model guide

Ministral 3 14B: Open-Weight Vision-Language AI for Local Deployment

Ministral 3 14B is Mistral AI’s largest Ministral 3 edge model, combining text generation, image understanding, function calling, structured outputs, and a 256k-token context window in an Apache 2.0 open-weight model designed for local and private deployment.

What is Ministral 3 14B?

Ministral 3 14B is Mistral AI’s largest model in the Ministral 3 family. It is an open-weight language and vision model designed for edge, private, and local deployments. The family includes Base, Instruct, and Reasoning variants. For hosted inference, Mistral identifies the model with the dated API name ministral-14b-2512.

The Instruct version is post-trained for conversations and instruction-following tasks. In practical terms, it is intended to answer questions, summarize information, analyze documents and images, generate code, and interact with external tools. Mistral positions it as a smaller alternative to larger general-purpose models, with a substantially lower deployment footprint than many frontier-scale systems.

“14B” refers approximately to the model’s 14 billion parameters. Parameter count is not a direct measure of quality, but it provides useful context for deployment: Ministral 3 14B is considerably more demanding than a small edge model while remaining feasible for local serving with suitable hardware.

Where it fits in Mistral AI’s lineup

Ministral 3 14B is part of Mistral AI’s compact, deployment-oriented Ministral 3 family. It is the largest member of that family and is aimed at users who need more capability than a very small edge model can provide without moving to a much larger model such as Mistral Small 3.2 24B.

The model is available in open-weight form under the Apache 2.0 license. That licensing approach gives organizations the option to download and operate the model in their own environment, subject to the license terms and their own operational responsibilities. It can also be accessed through Mistral’s hosted inference services.

The distinction between the family variants matters. Base checkpoints are intended for developers working closer to the model level, the Instruct checkpoint is optimized for following user directions, and the Reasoning variant is intended for tasks that benefit from more deliberate problem solving. The supplied research does not establish a separate maximum output limit for each variant.

Key capabilities and supported inputs

Ministral 3 14B accepts text and images as inputs and returns text. It can therefore support image-grounded conversations, document inspection, screenshot analysis, diagram interpretation, and ordinary text generation. It is multimodal on the input side, not a model that produces images, audio, video, or music.

  • Text generation: It can produce conversational answers, summaries, explanations, classifications, and other language-based outputs.
  • Image understanding: Images can be supplied with text prompts for visual analysis.
  • Function calling: The model can generate arguments for developer-defined tools or functions.
  • Structured output: It can produce constrained machine-readable responses through Mistral’s structured-output features.
  • Multilingual processing: The model card identifies languages including English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, and Arabic.
  • Long-context processing: The documented context window is 262,144 tokens, commonly described as 256k tokens.

Function calling does not mean that the model independently performs an external action. It generates a proposed tool call and its arguments; the surrounding application must validate those arguments, execute the tool, and return the result to the model. This makes it suitable for agentic applications, but the developer remains responsible for permissions, error handling, and safety controls.

Context window and output limits

The documented context length is 262,144 tokens. A context window is the total amount of text and other tokenized information that a request can contain, including the conversation history, instructions, tool definitions, retrieved documents, and the requested response. This large limit is useful for long documents, collections of notes, source code, and image-grounded conversations.

The full context capacity has practical hardware consequences for local serving. A deployment that technically supports 256k tokens may still need substantial memory for the model, the attention cache, image inputs, and concurrent users. Operators on constrained systems may reduce the maximum model length to improve memory usage and stability.

No authoritative maximum output-token limit was identified in the supplied model research. The 256k figure should therefore not be interpreted as a guaranteed response length: it describes the overall context capacity, not necessarily the maximum number of tokens that can be generated in one response.

Local hardware and deployment

Ministral 3 14B is specifically attractive when data locality or deployment control matters. Organizations can use the open-weight checkpoint in private infrastructure rather than sending every document or image to a third-party hosted endpoint. Potential applications include internal assistants, private document analysis, on-device or edge workflows, and systems that must operate within an organization’s own network.

Mistral’s deployment guidance covers tools including vLLM and Transformers. The FP8 Instruct checkpoint is documented as fitting in approximately 24 GB of VRAM. Quantization can reduce memory requirements, although the exact result depends on the quantization format, runtime, context length, batch size, and workload.

For vLLM deployments, Mistral documents model-specific tokenizer and configuration settings. Tool use requires automatic tool choice and the Mistral tool-call parser. These details are important because a model may load successfully while tool calls still fail if the serving engine is not configured to interpret its output correctly.

Local deployment is not automatically inexpensive or simple. A 14B model still requires meaningful GPU or system memory, and long-context requests can increase memory consumption substantially. Teams should test the intended quantization, context length, concurrency, and image workload rather than relying only on the parameter count.

Hosted pricing and model identifiers

Mistral’s hosted inference pricing lists Ministral 3 14B at $0.20 per million input tokens and $0.20 per million output tokens. Cached input is listed at $0.02 per million tokens. These are usage-based inference prices rather than a monthly subscription price for the model.

The dated identifier is ministral-14b-2512. Some Mistral tooling also exposes ministral-14b-latest, but that is a rolling pointer and should not be treated as a separate model. Applications that require reproducibility should prefer the dated identifier when it is supported by the relevant interface.

At this price, the model is positioned as a relatively cost-efficient option for hosted text and vision workloads. Actual spending depends on prompt size, output length, image tokenization, caching, request volume, and whether the application sends large conversation histories repeatedly.

Reasoning, coding, and tool use

Ministral 3 14B supports general reasoning and problem-solving, with a separate Reasoning variant available in the family. The supplied research does not provide a standardized benchmark result, so claims about its reasoning quality should be treated as practical positioning rather than a guarantee of performance on every task.

For coding, the model can generate and explain code, transform snippets, review files, and participate in tool-driven development workflows. Its function-calling support also allows it to work as part of an application that can search data, call services, or operate approved business tools. However, the model itself does not execute code or tools merely because it generated a call. Execution must be implemented and controlled by the host application.

Its combination of structured outputs and tool calling is useful when downstream software needs predictable fields rather than free-form prose. Developers should still validate generated JSON, enforce schemas, and treat tool arguments as untrusted input.

Main strengths and trade-offs

The most important strength of Ministral 3 14B is the balance between capability and deployability. It provides vision input, long context, multilingual support, function calling, and structured output while remaining small enough to consider for private serving. The Apache 2.0 open-weight release is also significant for teams that need more control over hosting, data handling, or customization than a closed hosted model provides.

Its cost and speed profile can be favorable compared with larger models, particularly for high-volume applications that do not need frontier-scale reasoning. The 14B size is still large enough to require careful hardware planning, however. Smaller models may be faster and easier to run, while larger models may provide stronger results on especially difficult reasoning, coding, or knowledge tasks.

The long context is useful but should not be confused with perfect comprehension. Sending more information increases memory and potentially cost, and the model may still miss relevant details in a very large prompt. Retrieval, document segmentation, clear instructions, and evaluation remain important.

Limitations to consider

  • It produces text and structured text, not native images, video, audio, speech, or music.
  • The supplied research does not verify a model-specific knowledge cutoff.
  • No authoritative maximum output-token value is provided.
  • Local serving requires meaningful memory, especially with the full context window, high concurrency, or image inputs.
  • Tool calls require an external application to execute and secure them.
  • The Base, Instruct, and Reasoning checkpoints are not interchangeable; each is intended for a different use.
  • Open weights provide deployment control but shift infrastructure, monitoring, security, and update responsibilities to the operator.

These limitations make Ministral 3 14B a poor fit for applications that require native media generation or a fully managed assistant with no infrastructure work. It may also be less suitable than a larger model when the primary requirement is the strongest available performance on difficult reasoning or code-generation tasks.

When to choose Ministral 3 14B

Choose Ministral 3 14B when you need a compact open-weight model that can understand both text and images, support long documents, call application tools, and run either through hosted inference or in private infrastructure. It is especially well suited to private assistants, multilingual document processing, screenshot and diagram analysis, internal enterprise applications, and cost-conscious agentic systems.

Choose a smaller model when low memory use, very fast responses, or inexpensive edge deployment matters more than the additional capability of a 14B model. Choose a larger model when your evaluations show that difficult reasoning, advanced coding, or complex multi-step tasks justify higher compute and inference costs. Choose a dedicated media-generation model when the required output is an image, video, audio clip, or speech file.

Overall, Ministral 3 14B is best understood as a flexible middle-ground model: more capable and multimodal than many lightweight local models, but more deployable and potentially less expensive than much larger frontier systems. Its strongest practical differentiator is not a single benchmark claim; it is the combination of open weights, image understanding, long context, tool support, and a hardware footprint that remains realistic for private deployment.


Answers to Frequently Asked Questions

How much does Ministral 3 14B cost through hosted inference?
Mistral lists Ministral 3 14B at $0.20 per million input tokens and $0.20 per million output tokens. Cached input is listed at $0.02 per million tokens. The dated hosted model identifier is ministral-14b-2512, while ministral-14b-latest is a rolling model pointer.
What is the context window of Ministral 3 14B?
Ministral 3 14B has a documented context window of 262,144 tokens, commonly described as 256k tokens. This total includes conversation history, instructions, tool definitions, retrieved documents, image-related inputs, and the generated response. It should not be interpreted as a guaranteed maximum output length.
How much hardware does Ministral 3 14B require for local deployment?
The documented FP8 Instruct checkpoint fits in approximately 24 GB of VRAM, although actual requirements vary by quantization format, context length, batch size, image inputs, runtime, and concurrency. Long-context requests can substantially increase memory use, so deployment should be tested with the intended workload.
What are the main capabilities of Ministral 3 14B?
Ministral 3 14B supports text generation, image understanding, multilingual processing, long-context tasks, function calling, and structured outputs. It can analyze documents, screenshots, and diagrams, generate or review code, summarize information, and interact with tools through a host application.
What is Ministral 3 14B?
Ministral 3 14B is Mistral AI’s largest Ministral 3 model, an open-weight language and vision model designed for local, private, edge, and hosted deployments. It accepts text and images, generates text, and is available in Base, Instruct, and Reasoning variants.


Sources 7
Provider

About Mistral AI