What is Mistral OCR 3?
Mistral OCR 3 is a specialized document-understanding model provided by Mistral AI. Its canonical API model identifier is mistral-ocr-2512. It accepts supported documents such as PDFs and images, then extracts their text and embedded images into machine-readable results.
Unlike a general-purpose language model, OCR 3 is focused on turning visually formatted documents into usable data. Its output can preserve document structure in markdown and can represent complex tables with HTML. That makes the model suitable for systems that need to search, index, classify, or further process the contents of business and archival documents.
The model was released on December 17, 2025. In Mistral AI's current catalog, OCR 4 is the newer OCR model, while OCR 3 remains available for existing integrations and production workloads. This makes OCR 3 particularly relevant when an application already depends on the mistral-ocr-2512 identifier or needs compatibility with an established OCR 3 deployment.
Where OCR 3 is strongest
OCR 3 is designed for documents that are difficult for basic text-recognition systems. Mistral describes improvements over the previous OCR generation for forms, handwritten content, low-quality scans, and complex tables.
- Forms: It can interpret printed forms and annotations placed over them, supporting workflows that need to connect labels, fields, and handwritten responses.
- Handwriting: It is intended to handle cursive writing and handwritten annotations, including writing that appears alongside printed material.
- Scanned documents: The model is designed for skewed pages, compression artifacts, low-resolution scans, and background noise that can reduce the accuracy of simpler OCR systems.
- Tables: It can reconstruct headers, merged cells, multi-row sections, and column hierarchies. HTML output can use elements such as
colspanandrowspanto preserve relationships between cells. - Embedded images: OCR results can include references to images contained in the source document, rather than treating the file as plain text only.
These capabilities are important because document meaning often depends on layout. A table's column hierarchy, a handwritten value beside a printed label, or an annotation over a form can be lost when a page is reduced to a simple sequence of words.
Inputs, outputs, and structured extraction
The documented input types are PDFs and images. OCR 3 extracts text and embedded images from those documents. Its primary output is document content, not generated artwork, speech, video, or conversational responses.
Standard results can be returned as structured markdown. Markdown is useful when the goal is to create searchable or readable document representations while retaining headings, paragraphs, lists, and other basic structure. For complex tables, HTML representation can preserve merged cells and more detailed relationships than plain text or uncomplicated markdown.
The OCR API also supports annotation workflows. These can request structured extraction as a JSON object or according to a JSON Schema. In practical terms, an application can use OCR 3 not only to transcribe a document but also to request fields such as an invoice number, date, supplier, or total in a predictable structure. The supplied research does not specify a maximum context length or maximum output-token limit, so those values should not be assumed when designing very large-document jobs.
Pricing and API access
Mistral OCR 3 is available through Mistral's OCR API and through the Document AI interface in Mistral AI Studio. Its documented standard price is $2 per 1,000 pages. The primary pricing unit is pages rather than input and output tokens, which makes high-volume document-processing costs relatively straightforward to estimate.
Annotation processing is documented at $3 per 1,000 annotated pages. Mistral's launch announcement also describes a 50% Batch API discount for standard processing, reducing the standard batch price to approximately $1 per 1,000 pages. These prices refer to the documented OCR processing modes; actual expenditure depends on the number of pages and whether annotation or batch processing is used.
The supplied research does not identify a separate subscription plan, free allowance, context limit, or output limit for OCR 3. Buyers should therefore verify current API billing details and operational limits in Mistral's documentation before estimating a production migration.
Capabilities and trade-offs
| Area | What is supported or known |
|---|---|
| Primary task | Document OCR and document understanding |
| Document inputs | PDFs and images |
| Text output | Yes, including structured markdown |
| Embedded images | Extracted or referenced in OCR results |
| Structured output | JSON object and JSON Schema annotation modes |
| Tables | Complex table reconstruction, including HTML relationships such as merged cells |
| Audio and video | Not supported as documented OCR inputs or outputs |
| Tool use | No general-purpose tool or function-calling capability is documented for this model |
| Fine-tuning | Not documented as supported in the supplied research |
| Streaming | Not documented as supported |
| Reasoning and coding | Not intended for open-ended reasoning or software development |
The research classifies OCR 3's reasoning capability as low for general reasoning and its coding capability as very low. Those are editorial evaluations rather than provider-published benchmark results. They reflect the model's narrow purpose: it can interpret document structure and extract information, but it should not be selected as a general reasoning or programming model.
Similarly, the supplied editorial scoring rates OCR 3 highly for speed and cost. These scores should not be treated as measured provider benchmarks. The concrete cost information is the page-based price published by Mistral, while real processing speed will depend on document size, workload, API conditions, and the chosen processing mode.
Practical use cases
OCR 3 is a good fit when the main challenge is converting visually complex documents into data that another system can use. Suitable applications include:
- Digitizing scanned business, historical, and archival records.
- Extracting invoice, receipt, and purchase-order fields.
- Processing insurance, compliance, legal, or administrative forms.
- Reading handwritten responses and annotations on printed documents.
- Converting technical and scientific PDFs into searchable markdown.
- Preserving complex tables for analytics or downstream data extraction.
- Preparing documents for retrieval-augmented generation, enterprise search, and AI-agent workflows.
For example, an accounts-payable pipeline could submit invoice pages to OCR 3, preserve the invoice's table structure, and request a JSON annotation containing the supplier, invoice number, line items, tax, and total. A records-management system could instead use the markdown output to index scanned documents while retaining references to embedded images.
Limitations and current status
OCR 3 is not a general-purpose language model. It is not the appropriate choice for ordinary conversation, autonomous agents, image generation, speech processing, video generation, or general coding assistance. Its output is extracted or structured document content, so applications should validate important fields before using them for financial, legal, medical, or compliance decisions.
The absence of a published context-length or maximum-output specification in the supplied research is also relevant. Teams processing very long documents should test page limits, request sizing, output behavior, and failure handling with their intended API workflow rather than assuming that an entire document can always be processed in one request.
OCR 3's current status is best described as compatibility-focused or legacy within Mistral's OCR lineup. Mistral identifies OCR 4 as the newer model, and the company has also announced OCR 4.1. New projects should compare the current OCR models against OCR 3 for accuracy, supported features, price, and migration requirements. Existing production systems may still prefer OCR 3 when stability around mistral-ocr-2512 matters more than moving immediately to a newer version.
When to choose Mistral OCR 3
Choose OCR 3 when you need page-based document extraction and your documents contain forms, handwriting, poor scans, or tables whose layout matters. It is especially reasonable when an existing integration already uses mistral-ocr-2512, when predictable per-page budgeting is useful, or when you need markdown, HTML table reconstruction, and structured JSON annotation in the same document workflow.
Choose a newer OCR option, including Mistral's OCR 4 generation, when starting a new implementation and the newer model offers better support for your documents or a more favorable current price and availability profile. Choose a general-purpose language model instead when the central requirement is conversation, code generation, broad reasoning, or tool-driven task execution rather than document extraction.
In short, OCR 3 should be evaluated as a focused document-processing component. Its value comes from recognizing difficult page content and preserving useful structure, not from the broad interactive capabilities associated with general AI models.

