OCR

OCR 3

by Mistral AI · Legacy; available for existing integrations and production workloads. OCR 4 is the newer model.

A specialized Mistral AI OCR model for extracting text and embedded images from PDFs and images, with particular support for handwriting, forms, poor-quality scans, complex tables, markdown, HTML table reconstruction, and JSON-based document annotation.

Text Image generation Reasoning Coding
Mistral OCR 3 is a production-oriented OCR and document-understanding service released in December 2025. It converts PDFs and images into structured markdown and can preserve complex table relationships with HTML. The model is intended for digitization, document search, invoice and form processing, compliance workflows, and retrieval-augmented generation rather than general chat or coding.
Outputs

What OCR 3 can produce

Text Image generation
Inputs

What it can understand

Images Multimodal input
Capabilities

Supported features

JSON mode Structured output Batch API Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family OCR
Model type Other
Release date 2025-12-17
Status Legacy; available for existing integrations and production workloads. OCR 4 is the newer model.
Knowledge cutoff notes

Mistral does not publish a conventional knowledge-cutoff date for OCR 3. It is a document extraction model whose behavior depends on the supplied document rather than a stated training-data cutoff.

Model notes

The canonical API identifier is mistral-ocr-2512. OCR 3 extracts text and embedded images from supported documents and can reconstruct complex tables using HTML. The standard API price is page-based rather than token-based. Mistral documents OCR 4 as the newer model but states that OCR 3 remains available for existing integrations and production workloads. The model supports OCR output in text or structured annotation formats, including JSON object and JSON schema modes. Batch processing is documented with a 50% discount in Mistral's launch announcement, reducing the standard page price to approximately $1 per 1,000 pages.

Cost

Model pricing

Input $2 per 1,000 pages
Output $3 per 1,000 annotated pages
Model guide

Mistral OCR 3 for Handwriting, Forms, and Complex Tables

Mistral OCR 3 is a specialized Mistral AI document-understanding model for extracting text and embedded images from PDFs and images. Its canonical API identifier is mistral-ocr-2512, and it is designed for forms, handwriting, scanned documents, invoices, and complex tables. OCR 3 remains available for existing integrations and production workloads, although OCR 4 is the newer model.

What is Mistral OCR 3?

Mistral OCR 3 is a specialized document-understanding model provided by Mistral AI. Its canonical API model identifier is mistral-ocr-2512. It accepts supported documents such as PDFs and images, then extracts their text and embedded images into machine-readable results.

Unlike a general-purpose language model, OCR 3 is focused on turning visually formatted documents into usable data. Its output can preserve document structure in markdown and can represent complex tables with HTML. That makes the model suitable for systems that need to search, index, classify, or further process the contents of business and archival documents.

The model was released on December 17, 2025. In Mistral AI's current catalog, OCR 4 is the newer OCR model, while OCR 3 remains available for existing integrations and production workloads. This makes OCR 3 particularly relevant when an application already depends on the mistral-ocr-2512 identifier or needs compatibility with an established OCR 3 deployment.

Where OCR 3 is strongest

OCR 3 is designed for documents that are difficult for basic text-recognition systems. Mistral describes improvements over the previous OCR generation for forms, handwritten content, low-quality scans, and complex tables.

  • Forms: It can interpret printed forms and annotations placed over them, supporting workflows that need to connect labels, fields, and handwritten responses.
  • Handwriting: It is intended to handle cursive writing and handwritten annotations, including writing that appears alongside printed material.
  • Scanned documents: The model is designed for skewed pages, compression artifacts, low-resolution scans, and background noise that can reduce the accuracy of simpler OCR systems.
  • Tables: It can reconstruct headers, merged cells, multi-row sections, and column hierarchies. HTML output can use elements such as colspan and rowspan to preserve relationships between cells.
  • Embedded images: OCR results can include references to images contained in the source document, rather than treating the file as plain text only.

These capabilities are important because document meaning often depends on layout. A table's column hierarchy, a handwritten value beside a printed label, or an annotation over a form can be lost when a page is reduced to a simple sequence of words.

Inputs, outputs, and structured extraction

The documented input types are PDFs and images. OCR 3 extracts text and embedded images from those documents. Its primary output is document content, not generated artwork, speech, video, or conversational responses.

Standard results can be returned as structured markdown. Markdown is useful when the goal is to create searchable or readable document representations while retaining headings, paragraphs, lists, and other basic structure. For complex tables, HTML representation can preserve merged cells and more detailed relationships than plain text or uncomplicated markdown.

The OCR API also supports annotation workflows. These can request structured extraction as a JSON object or according to a JSON Schema. In practical terms, an application can use OCR 3 not only to transcribe a document but also to request fields such as an invoice number, date, supplier, or total in a predictable structure. The supplied research does not specify a maximum context length or maximum output-token limit, so those values should not be assumed when designing very large-document jobs.

Pricing and API access

Mistral OCR 3 is available through Mistral's OCR API and through the Document AI interface in Mistral AI Studio. Its documented standard price is $2 per 1,000 pages. The primary pricing unit is pages rather than input and output tokens, which makes high-volume document-processing costs relatively straightforward to estimate.

Annotation processing is documented at $3 per 1,000 annotated pages. Mistral's launch announcement also describes a 50% Batch API discount for standard processing, reducing the standard batch price to approximately $1 per 1,000 pages. These prices refer to the documented OCR processing modes; actual expenditure depends on the number of pages and whether annotation or batch processing is used.

The supplied research does not identify a separate subscription plan, free allowance, context limit, or output limit for OCR 3. Buyers should therefore verify current API billing details and operational limits in Mistral's documentation before estimating a production migration.

Capabilities and trade-offs

AreaWhat is supported or known
Primary taskDocument OCR and document understanding
Document inputsPDFs and images
Text outputYes, including structured markdown
Embedded imagesExtracted or referenced in OCR results
Structured outputJSON object and JSON Schema annotation modes
TablesComplex table reconstruction, including HTML relationships such as merged cells
Audio and videoNot supported as documented OCR inputs or outputs
Tool useNo general-purpose tool or function-calling capability is documented for this model
Fine-tuningNot documented as supported in the supplied research
StreamingNot documented as supported
Reasoning and codingNot intended for open-ended reasoning or software development

The research classifies OCR 3's reasoning capability as low for general reasoning and its coding capability as very low. Those are editorial evaluations rather than provider-published benchmark results. They reflect the model's narrow purpose: it can interpret document structure and extract information, but it should not be selected as a general reasoning or programming model.

Similarly, the supplied editorial scoring rates OCR 3 highly for speed and cost. These scores should not be treated as measured provider benchmarks. The concrete cost information is the page-based price published by Mistral, while real processing speed will depend on document size, workload, API conditions, and the chosen processing mode.

OCR 3 is a good fit when the main challenge is converting visually complex documents into data that another system can use. Suitable applications include:

  • Digitizing scanned business, historical, and archival records.
  • Extracting invoice, receipt, and purchase-order fields.
  • Processing insurance, compliance, legal, or administrative forms.
  • Reading handwritten responses and annotations on printed documents.
  • Converting technical and scientific PDFs into searchable markdown.
  • Preserving complex tables for analytics or downstream data extraction.
  • Preparing documents for retrieval-augmented generation, enterprise search, and AI-agent workflows.

For example, an accounts-payable pipeline could submit invoice pages to OCR 3, preserve the invoice's table structure, and request a JSON annotation containing the supplier, invoice number, line items, tax, and total. A records-management system could instead use the markdown output to index scanned documents while retaining references to embedded images.

Limitations and current status

OCR 3 is not a general-purpose language model. It is not the appropriate choice for ordinary conversation, autonomous agents, image generation, speech processing, video generation, or general coding assistance. Its output is extracted or structured document content, so applications should validate important fields before using them for financial, legal, medical, or compliance decisions.

The absence of a published context-length or maximum-output specification in the supplied research is also relevant. Teams processing very long documents should test page limits, request sizing, output behavior, and failure handling with their intended API workflow rather than assuming that an entire document can always be processed in one request.

OCR 3's current status is best described as compatibility-focused or legacy within Mistral's OCR lineup. Mistral identifies OCR 4 as the newer model, and the company has also announced OCR 4.1. New projects should compare the current OCR models against OCR 3 for accuracy, supported features, price, and migration requirements. Existing production systems may still prefer OCR 3 when stability around mistral-ocr-2512 matters more than moving immediately to a newer version.

When to choose Mistral OCR 3

Choose OCR 3 when you need page-based document extraction and your documents contain forms, handwriting, poor scans, or tables whose layout matters. It is especially reasonable when an existing integration already uses mistral-ocr-2512, when predictable per-page budgeting is useful, or when you need markdown, HTML table reconstruction, and structured JSON annotation in the same document workflow.

Choose a newer OCR option, including Mistral's OCR 4 generation, when starting a new implementation and the newer model offers better support for your documents or a more favorable current price and availability profile. Choose a general-purpose language model instead when the central requirement is conversation, code generation, broad reasoning, or tool-driven task execution rather than document extraction.

In short, OCR 3 should be evaluated as a focused document-processing component. Its value comes from recognizing difficult page content and preserving useful structure, not from the broad interactive capabilities associated with general AI models.


Answers to Frequently Asked Questions

When should you choose Mistral OCR 3 instead of a newer OCR model?
Choose Mistral OCR 3 when you need to process forms, handwriting, poor scans, or structurally complex tables, especially if an existing integration depends on mistral-ocr-2512. For new projects, compare it with newer models such as Mistral OCR 4 for accuracy, features, price, availability, and migration requirements.
What does Mistral OCR 3 cost?
The documented standard price is $2 per 1,000 pages. Annotation processing costs $3 per 1,000 annotated pages, while the Batch API discount described by Mistral reduces standard batch processing to approximately $1 per 1,000 pages.
How does Mistral OCR 3 handle complex tables?
Mistral OCR 3 can reconstruct complex table structures, including headers, merged cells, multi-row sections, and column hierarchies. Its HTML output can preserve relationships using elements such as colspan and rowspan.
What is Mistral OCR 3 and what model identifier does it use?
Mistral OCR 3 is a document-understanding and optical character recognition model for extracting text, structure, and embedded images from PDFs and images. Its canonical API model identifier is mistral-ocr-2512.
Can Mistral OCR 3 read handwriting, forms, and low-quality scans?
Yes. Mistral OCR 3 is designed to process handwritten content, cursive writing, annotations on printed forms, skewed pages, compression artifacts, low-resolution scans, and noisy backgrounds.


Sources 5
Provider

About Mistral AI