What is Mistral Large 3?
Mistral Large 3 is an open-weight multimodal model from Mistral AI. It was released on December 2, 2025, and is listed by Mistral as an active, generally available model. Its primary role is general-purpose language and vision processing: it can work with ordinary text, analyze images, follow instructions, generate code, call external tools, and return structured responses.
The model is available in two main forms. Organizations can use it through Mistral AI's commercial API, or download its weights for self-hosting, customization, and more controlled deployment. Both the Base and instruction-tuned versions are distributed under the Apache 2.0 license according to the supplied model information. The canonical API model ID is mistral-large-2512, while mistral-large-latest is Mistral's rolling API alias.
Mistral Large 3 belongs to the Mistral Large family but is also part of Mistral's broader open-weight model strategy. Its combination of a very large total parameter count, sparse activation, long context, image understanding, and deployment flexibility positions it toward demanding enterprise and research workloads rather than lightweight local use.
Architecture and capacity
Mistral Large 3 uses a granular mixture-of-experts, or MoE, architecture. In an MoE model, different portions of the network are selected for different inputs instead of activating every parameter for every token. Mistral Large 3 has 675 billion total parameters and approximately 41 billion active parameters. The active count helps explain why the model can offer sparse inference compared with a dense model containing the same total number of parameters, but it remains a very large system to operate.
The documented context window is 256,000 tokens. Context is the amount of text and other supported input that the model can consider in a request. A 256K window is useful for large documents, multi-file analysis, long conversations, extensive retrieval results, and codebases that would otherwise need to be split into many smaller requests. The context limit is not the same as a guaranteed output length: the supplied documentation does not identify a separate maximum output-token limit for Mistral Large 3.
The large context window does not remove the need for application design. Very large prompts can increase processing time and cost, and long inputs still need careful retrieval, ordering, and validation. For repeated prefixes or shared instructions, Mistral documents prefix or prompt caching, which can reduce the cost of cached input tokens.
Supported inputs and outputs
Mistral Large 3 accepts text and images. Its vision encoder allows it to interpret image inputs alongside language, making it suitable for image-aware document analysis, screenshots, diagrams, and other workflows where visual information needs to be combined with written instructions.
The model produces text output. It does not natively generate images, audio, or video. This distinction matters because “multimodal” describes the model's ability to process more than one input type; it does not mean that the model can generate every kind of media. Image creation or speech generation would require a separate compatible service or model.
For application integration, Mistral documents structured outputs and function calling. Structured outputs are useful when a program needs predictable fields such as a JSON object, classification result, extracted invoice data, or a list of entities. Function calling allows the model to request an application-defined operation, such as searching a database, retrieving a document, or initiating a workflow. The application remains responsible for executing the function and checking its arguments.
Reasoning, coding, and tool use
Mistral Large 3 is intended for complex knowledge work, including research, coding, retrieval-augmented generation, document question answering, and agentic workflows. Its editorial reasoning and coding assessments are both 8 out of 10; these are comparative editorial evaluations, not provider-published benchmark scores. They indicate that the model is positioned as a capable general-purpose system for multi-step tasks, but they should not be treated as measured guarantees.
For coding, the model can generate and explain code, analyze programming materials, and participate in tool-driven development workflows. Coding quality will depend on the language, repository context, instructions, and validation process. A long context can help when supplying several related files, while structured outputs and function calling can support coding agents that need to return machine-readable plans or invoke development tools.
Mistral documents chat completions, streaming, function calling, structured outputs, document question answering, batching, and the Agents and Conversations interfaces. Built-in web search is available through the Conversations and Agents APIs for supported models, including Mistral Large 3 according to the supplied research. Web search provides a retrieval mechanism during use; it does not establish a fixed or current underlying training-data cutoff. No authoritative knowledge-cutoff date was identified for this exact model.
API pricing and deployment options
Current Mistral AI Studio pricing for Mistral Large 3 is $0.50 per million input tokens, $0.05 per million cached input tokens, and $1.50 per million output tokens. Input and output are billed separately, so applications that produce long responses or repeatedly send large prompts should estimate both sides of usage. Cached-input pricing applies to eligible repeated content and should not be confused with the standard input rate.
The model's open weights provide an alternative to API-only use. Mistral highlights FP8 and NVFP4 deployment options for high-memory GPU systems, including configurations based on H100, H200, B200, and A100 hardware. These options may be relevant to organizations requiring private, sovereign, or customized deployments, but the model's 675-billion total parameter size means self-hosting demands substantial accelerator memory, storage, networking, and operational expertise.
API access is therefore likely to be simpler for teams that need fast integration without managing infrastructure. Downloadable weights are more relevant when deployment control, data locality, fine-tuning, or operational independence outweighs the complexity and capital cost of running the model. The editorial cost assessment is 8 out of 10, reflecting the published token pricing and open-weight flexibility rather than a guarantee that the model will be cheaper than every alternative.
Main strengths and limitations
Key strengths
- Open deployment model: Apache 2.0 weights and commercial API access give organizations more choice than a service available only through a hosted endpoint.
- Long context: The 256,000-token context window supports large documents, extended conversations, code repositories, and retrieval-heavy prompts.
- Image understanding: Text and image inputs can be processed together for visual document and screenshot analysis.
- Application integration: Function calling, structured outputs, streaming, batching, and Agents and Conversations interfaces support production workflows.
- Multilingual and general-purpose use: The model is designed for multilingual interaction, document analysis, coding, research, and enterprise assistants.
- Deployment flexibility: API access and downloadable weights accommodate both managed and self-hosted architectures.
Important limitations
- Infrastructure demands: Although MoE sparsity reduces the number of active parameters per token, 675 billion total parameters still makes local deployment unsuitable for ordinary consumer hardware.
- No native media generation: The documented multimodality covers image understanding and text generation, not image, audio, or video creation.
- Output limit uncertainty: A separate maximum output-token limit was not identified in the current model documentation.
- Variable cost at scale: Long prompts and long responses can increase API bills even when the per-million-token rates appear manageable.
- Probabilistic results: Function arguments, extracted data, code, and generated explanations require application-level validation.
- Self-hosting complexity: Downloadable weights do not make deployment effortless; suitable GPUs and production infrastructure remain necessary.
When to choose Mistral Large 3
Choose Mistral Large 3 when you need a single general-purpose model that combines long-context text processing, image understanding, tool use, and open deployment options. It is a strong candidate for enterprise assistants that must read substantial internal material, multilingual document-analysis systems, retrieval-augmented applications, coding and research agents, and organizations that want the option of private infrastructure.
It is particularly suitable when control over model weights or deployment location is important. A team can begin with the API for faster implementation and evaluate self-hosting later if data-governance, customization, or infrastructure requirements justify the transition. The Apache 2.0 licensing and available Base and Instruct weights may also be important for organizations building products around a modifiable model.
Another option may be more appropriate when the workload requires native image, audio, or video generation, because Mistral Large 3 is a text-output model with image understanding rather than a general media-generation system. A smaller model may be preferable for low-latency, high-volume, or low-memory workloads. Conversely, an API-only model may be easier for teams that cannot support the hardware and operations required by a 675-billion-parameter open-weight system.
The model's speed assessment is an editorial 5 out of 10, not a published latency benchmark. This reflects the practical trade-off between broad capability and the resources required by a model of this scale. Users should test their own prompts, hardware, concurrency, and response-length requirements before committing to a production architecture.
Bottom line
Mistral Large 3 is designed for demanding text-and-image workloads where long context, tool integration, and deployment flexibility matter more than minimal infrastructure requirements. Its verified specifications include 675 billion total parameters, approximately 41 billion active parameters, a 256,000-token context window, text and image input, text output, function calling, structured outputs, batching, and API pricing of $0.50 per million input tokens and $1.50 per million output tokens, with lower pricing for cached input.
Its most important practical distinction is the combination of open weights and a hosted API. That makes it more adaptable than a purely closed, hosted model, but not automatically easier or cheaper to operate. Teams should select it for long-context, multilingual, image-aware, agentic, or self-hosted applications, and choose a smaller or media-generation-focused option when those are the dominant requirements.

