Mistral OCR

OCR 4.0

by Mistral AI · Generally available; superseded by OCR 4.1 as Mistral's latest OCR model

Mistral OCR 4.0 is a specialized document-understanding model for extracting text and layout information from PDFs, images, presentations, and office documents. It supports structural block labels, paragraph-level bounding boxes, confidence scores, markdown, and JSON or JSON-schema-constrained annotations. Standard API pricing is $4 per 1,000 pages, with separate cached and Batch API rates. OCR 4.1 is now the newer model in Mistral's OCR family.

Text Reasoning Coding
Mistral OCR 4.0 goes beyond turning scanned pages into plain text. It identifies document elements such as titles, paragraphs, lists, tables, images, equations, headers, footers, captions, and signatures, while also reporting where those elements appear on the page. That combination makes it useful when an application needs to search, cite, classify, validate, or automate work with real-world documents rather than merely display a transcription.
Outputs

What OCR 4.0 can produce

Text
Inputs

What it can understand

Images Multimodal input
Capabilities

Supported features

JSON mode Structured output Prompt caching Batch API
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Mistral OCR
Model type Other
Release date 2026-06-23
Status Generally available; superseded by OCR 4.1 as Mistral's latest OCR model
Knowledge cutoff notes

Mistral does not publish a separate knowledge-cutoff date for OCR 4.0. The model performs document extraction from supplied inputs rather than relying on a conventional public web-training cutoff.

Model notes

The exact canonical API model ID is mistral-ocr-4-0. OCR 4.0 accepts documents and images and returns extracted text with document structure. Its OCR endpoint supports JSON object mode and JSON schema mode, as well as block extraction, bounding boxes, and confidence-score granularity. OCR 4.1 became generally available on August 31, 2026 and is the newer model; OCR 4.0 remains listed in Mistral's model catalog and pricing documentation. Mistral's launch material describes approximately 170 supported languages and a self-hosting option for enterprise customers. Context-window and maximum-output-token values are not published as standard specifications for this page-based OCR service.

Cost

Model pricing

Input $4 per 1,000 pages; $0.40 per 1,000 cached pages; Batch API pricing reported at $2 per 1,000 pages
Output Included in page-based processing price; no separate output-token price
Model guide

Mistral OCR 4.0: Structured Document Extraction with Bounding Boxes

Mistral OCR 4.0 is a document-understanding model that extracts text from PDFs, images, presentations, office documents, and other supported files while preserving document structure. Its output can include paragraph-level bounding boxes, block classifications, confidence scores, markdown, and JSON or JSON-schema-constrained annotations. The model is designed for high-volume OCR, enterprise search, retrieval-augmented generation, invoice processing, compliance workflows, and document automation. It was released on June 23, 2026, costs $4 per 1,000 pages through the standard API, and remains available even though Mistral OCR 4.1 is now the newer model in the family.

What is Mistral OCR 4.0?

Mistral OCR 4.0 is a specialized document-intelligence model from Mistral AI. Its canonical API model identifier is mistral-ocr-4-0. The service accepts documents and images, extracts their textual content, and returns information about the document's layout and structure.

This is an OCR model rather than a general conversational model. OCR, or optical character recognition, converts text contained in scans, photographs, PDFs, and other visual documents into machine-readable data. Mistral OCR 4.0 adds document understanding to that process: instead of returning only a continuous transcription, it can indicate what each content block represents and where it is located.

Mistral released OCR 4.0 on June 23, 2026. It is generally available through Mistral's API and Document AI offerings. OCR 4.1 is now the newer model in Mistral's OCR lineup, but OCR 4.0 remains listed as an available model and retains its documented pricing and capabilities.

Document processing capabilities

OCR 4.0 returns extracted text in reading order and can preserve useful page structure. Supported structural categories include text, titles, lists, tables, images, equations, captions, code, references, side text, headers, footers, and signatures.

For example, an invoice-processing application could distinguish the invoice title from line-item tables and totals. A legal-document pipeline could separate headings, body paragraphs, footnotes, and signatures. A search system could use block positions to show a result in its original page context instead of treating the source as an undifferentiated text file.

The model is intended for PDFs, images, presentations, word-processing documents, and OpenDocument files. Mistral's launch material describes support for approximately 170 languages across 10 language groups. This is a provider-reported coverage claim; the quality of extraction can still vary with language, scan quality, typography, layout complexity, and the source document itself.

Layout information and bounding boxes

A bounding box is a set of coordinates describing where an element appears on a page. OCR 4.0 provides paragraph-level bounding boxes and can expose structural information for individual document blocks. This allows downstream software to connect extracted text with its original visual location.

Bounding boxes are useful for citation-aware retrieval, visual document viewers, redaction, table handling, semantic chunking, and human review. A validation interface, for instance, can highlight the exact region from which an extracted invoice number was read. This is a more actionable output than plain text alone.

Structured annotations and confidence information

The OCR endpoint can return ordinary extracted text or structured annotations. Developers can request JSON objects or JSON-schema-constrained output for document-specific extraction tasks. A schema might define fields for an invoice number, supplier, date, currency, and total, allowing the returned information to fit a downstream application more predictably.

OCR 4.0 also supports confidence information at page, block, or word granularity. Confidence scores do not guarantee that an extraction is correct, but they can help applications identify uncertain regions and route them to human review. This is particularly relevant for documents where a single character error could change a payment amount, account number, legal clause, or compliance result.

Inputs, outputs, and technical limits

OCR 4.0 supports text extraction from visual and document inputs, including PDFs, images, presentations, word-processing files, and OpenDocument files. Its direct output is text and structured document data, not generated images, audio, video, or speech. The model does not provide general-purpose text generation, image generation, or conversational reasoning as its primary function.

The documented OCR response can include page-level results, markdown, extracted images, structural blocks, annotations, and usage information. The exact returned fields depend on the request and endpoint options.

Mistral does not publish a conventional context-window or maximum-output-token specification for OCR 4.0. That is consistent with its page-based document-processing design: usage and pricing are measured primarily by pages rather than by an advertised text-token context limit. Applications should therefore test the size and format of their documents against the current API behavior instead of assuming a general language-model context window.

The research does not verify streaming support, tool or function calling, or fine-tuning for OCR 4.0. It is best understood as a document extraction endpoint rather than an agent model that independently calls tools or performs multi-step reasoning.

Pricing and access

Standard API pricing is $4 per 1,000 pages. Mistral's pricing documentation also lists cached-input pricing of $0.40 per 1,000 pages. The launch announcement reports Batch API processing at $2 per 1,000 pages, which can be relevant for asynchronous, high-volume workloads.

Mistral's Document AI no-code experience is priced separately at $5 per 1,000 pages. That price should not be confused with the standard OCR 4.0 API rate. Enterprise self-hosting is also described as an option for customers with data-residency, sovereignty, or privacy requirements, but availability and commercial terms may depend on an enterprise agreement rather than the public API price.

Because the service is priced by page, cost planning differs from token-priced language models. A document with sparse text and a document packed with small print may each count as pages even though their amount of extracted text differs substantially. Teams should estimate page volume, cached usage, batch eligibility, and any separate Document AI or enterprise costs before deployment.

Strengths and limitations

Where OCR 4.0 is strong

  • Structured extraction: It can preserve document blocks such as headings, lists, tables, equations, captions, and signatures rather than returning only a flat transcription.
  • Spatial information: Paragraph-level bounding boxes connect extracted content to its location on the page.
  • Downstream automation: JSON and JSON Schema support makes it suitable for extracting fields into business systems.
  • Quality review: Page-, block-, and word-level confidence information can help prioritize manual verification.
  • Multilingual document coverage: Mistral describes support for approximately 170 languages and 10 language groups.
  • High-volume economics: Page-based pricing, caching, and Batch API processing provide options for large ingestion workloads.

Important limitations

  • Not a general AI assistant: OCR 4.0 is not intended for open-ended conversation, coding assistance, image generation, or general reasoning.
  • No published context or output-token limits: Mistral does not provide standard context-window and maximum-output-token values for this page-based service.
  • Extraction still requires validation: Confidence scores can identify uncertainty but do not eliminate recognition errors.
  • Limited verified model features: Tool use, streaming, and fine-tuning are not verified in the supplied specifications.
  • Model succession: OCR 4.1 is the newer Mistral OCR model, so new integrations may need to compare its capabilities and migration implications with OCR 4.0.

These limitations matter most in high-stakes workflows. OCR 4.0 should not be used as the sole authority for medical diagnosis, legal conclusions, financial decisions, or safety-critical automation. Extracted information should be checked when an error could cause material harm.

Best use cases for Mistral OCR 4.0

OCR 4.0 is a good fit when an organization needs to ingest many documents and retain enough structure for later processing. Practical uses include:

  • Enterprise search: Index text while retaining page locations, headings, tables, and other document structure.
  • Retrieval-augmented generation: Create more meaningful chunks for a question-answering system and preserve source locations for citations.
  • Invoice and form processing: Extract selected fields from semi-structured documents into accounting or workflow systems.
  • Compliance review: Identify relevant sections, signatures, headers, footers, and uncertain passages for review.
  • Knowledge-base construction: Convert document collections into searchable, structured records.
  • Document automation: Feed page-aware OCR results into classification, validation, redaction, or approval workflows.

It is especially appropriate when the location and type of content matter as much as the words themselves. A basic text-only OCR system may be cheaper or simpler for straightforward transcription, while a general multimodal language model may be more appropriate when the task requires broad reasoning over a small number of documents. OCR 4.0's main trade-off is specialization: it offers document-oriented structure and page-based processing rather than a broad conversational feature set.

When to choose Mistral OCR 4.0

Choose OCR 4.0 when your primary problem is converting a large or varied document collection into structured, searchable data. Its bounding boxes, block labels, confidence information, and schema-constrained annotations are more directly useful than plain text extraction for enterprise workflows.

It may be a sensible option when predictable page-based pricing, Batch API processing, multilingual coverage, or an enterprise deployment path is important. The standard $4-per-1,000-page rate can also be attractive for high-volume ingestion, although actual cost depends on page count, caching, batching, and the chosen Mistral product.

Consider OCR 4.1 instead when starting a new integration and the latest model capabilities are more important than keeping an OCR 4.0 implementation. Consider a general multimodal model when the central task is reasoning, summarization, question answering, or interactive analysis rather than systematic OCR across a document collection. Consider a simpler OCR tool when you only need a basic transcription and do not need structural labels, bounding boxes, confidence granularity, or JSON extraction.

Overall assessment

Mistral OCR 4.0 is best evaluated as a structured document-ingestion component, not as a general-purpose AI model. Its distinguishing value is the combination of extracted text, document structure, page coordinates, confidence information, and structured annotations. Those outputs can reduce the work required to build search, retrieval, extraction, and review systems around PDFs and other complex documents.

The main decision is whether those capabilities justify using a specialized OCR service when a newer OCR model or a broader multimodal option may be available. For existing workflows that need OCR 4.0's documented behavior and page-based pricing, it remains a practical choice. For new deployments, teams should compare it directly with OCR 4.1 and validate representative documents before committing to a production migration or high-stakes automated process.


Answers to Frequently Asked Questions

Should developers choose Mistral OCR 4.0 or OCR 4.1?
OCR 4.1 is the newer model in Mistral's OCR lineup, so new integrations should compare its capabilities and migration requirements with OCR 4.0. OCR 4.0 may still be appropriate for existing workflows that depend on its documented behavior, structured extraction, bounding boxes, confidence information, or page-based pricing.
How much does Mistral OCR 4.0 cost?
Standard API pricing is $4 per 1,000 pages, while cached input is listed at $0.40 per 1,000 pages. Batch API processing is reported at $2 per 1,000 pages. Mistral's no-code Document AI offering is priced separately at $5 per 1,000 pages.
Can Mistral OCR 4.0 return structured JSON and confidence scores?
Yes. The OCR endpoint can return structured annotations as JSON or JSON Schema-constrained output for tasks such as extracting invoice numbers, suppliers, dates, currencies, and totals. It also supports confidence information at page, block, or word level to help identify content that may require manual review.
Does Mistral OCR 4.0 provide bounding boxes and document layout information?
Yes. Mistral OCR 4.0 provides paragraph-level bounding boxes and structural information for document blocks, helping applications connect extracted text to its original page location. This supports citation-aware retrieval, visual document viewers, redaction, table processing, semantic chunking, and human validation.
What is Mistral OCR 4.0 used for?
Mistral OCR 4.0 is a specialized document-intelligence model for extracting text and structure from PDFs, images, presentations, word-processing files, and OpenDocument files. It is designed for workflows such as enterprise search, invoice processing, compliance review, knowledge-base creation, and document automation.


Sources 5
Provider

About Mistral AI