OCR

OCR 4.1

by Mistral AI · Generally Available

Mistral OCR 4.1 is Mistral AI's generally available OCR and document-understanding service. It processes PDFs and images, extracts structured text and tables, identifies document blocks, returns paragraph-level bounding boxes and confidence scores, and supports batch processing with page-based pricing.

Text Reasoning Coding
Mistral OCR 4.1 is a specialized document-understanding model from Mistral AI. It converts supported PDFs and images into machine-readable text and document structure, including paragraphs, tables, titles, lists, equations, images, headers, footers, signatures, bounding boxes, and confidence scores. The model is available through Mistral's OCR API and Document AI tooling under the identifier mistral-ocr-4-1, with pricing based on pages rather than language-model tokens.
Outputs

What OCR 4.1 can produce

Text
Inputs

What it can understand

Images Multimodal input
Capabilities

Supported features

JSON mode Structured output Prompt caching Batch API
Model profile

Performance characteristics

3/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family OCR
Model type Other
Release date 2026-07-16
Status Generally Available
Model notes

OCR 4.1 is a specialized OCR service rather than a conventional token-based language model. The canonical model ID is mistral-ocr-4-1. Mistral-ocr-latest and mistral-ocr-4 point to OCR 4.1. The API accepts document and image inputs and can return extracted text, Markdown, tables, paragraph-level bounding boxes, structural block labels, and confidence scores. Standard API pricing is $4 per 1,000 pages; Mistral separately lists Document AI pricing at $5 per 1,000 annotated pages. Mistral's OCR launch materials state that batch processing provides a 50% discount, reducing the API price to $2 per 1,000 pages. Editorial scores are comparative estimates for this specialized OCR model and are not vendor benchmarks.

Cost

Model pricing

Input $4 per 1,000 pages; cached input $0.40 per 1,000 pages
Model guide

Mistral OCR 4.1: Structured Document Extraction at Page-Based Pricing

Mistral OCR 4.1 is Mistral AI's specialized OCR and document-understanding service for extracting text, layout structure, tables, bounding boxes, labels, and confidence scores from PDFs and images. It is designed for document-processing pipelines rather than conversation, reasoning, or code generation.

What is Mistral OCR 4.1?

Mistral OCR 4.1 is Mistral AI's current specialized optical character recognition (OCR) and document-understanding service. OCR means converting text contained in documents or images into data that software can search, index, analyze, or pass to another system. OCR 4.1 goes beyond returning one plain text stream: it is designed to preserve important parts of a document's structure.

The model accepts document and image inputs and can return extracted text, Markdown, tables, paragraph-level bounding boxes, structural block labels, and confidence scores. A bounding box identifies where an extracted element appears on the page. Structural labels indicate what a block represents, such as a title, paragraph, table, list, equation, image, caption, code block, reference, header, footer, or signature.

Its canonical API identifier is mistral-ocr-4-1. According to Mistral's changelog, mistral-ocr-latest and mistral-ocr-4 point to OCR 4.1. Mistral released the model on July 16, 2026 and marked it Generally Available on August 31, 2026. It is listed as the latest OCR model in Mistral's current catalog.

What information does it extract?

OCR 4.1 is useful when the location and role of text matter as much as the words themselves. For example, a document ingestion system can treat a heading differently from a footer, preserve a table as a table, or identify a signature block for later review.

  • Text: Extracted content from supported PDFs and images.
  • Layout: Paragraph-level bounding boxes and page-level organization.
  • Structural blocks: Labels for text, titles, lists, tables, images, equations, captions, code, references, headers, footers, and signatures.
  • Tables: Table extraction with supported Markdown or HTML formatting options.
  • Confidence information: Confidence scores can be returned at page, block, or word granularity.
  • Structured annotations: Machine-readable output for downstream document-processing systems.

These features make OCR 4.1 more suitable for structured extraction than a basic scanner-style OCR tool that only produces unformatted text. The exact usefulness of the output depends on the document quality, layout, language, and the way the consuming application validates the extracted data.

Where OCR 4.1 fits in Mistral AI's lineup

OCR 4.1 is a specialized model in Mistral AI's Document AI and API ecosystem. It is not a general-purpose conversational model, coding model, or autonomous agent. Its purpose is to transform visual documents into text and structured document data that can then be searched, stored, reviewed, or supplied to another model.

OCR 4.0 and OCR 3 remain separate catalog entries, but OCR 4.1 is the current canonical version for new integrations according to the supplied Mistral documentation. The service is therefore best understood as a document-ingestion component rather than a replacement for a general Mistral language model.

Inputs, outputs, and technical capabilities

AreaVerified information
Input typesDocuments and images, including PDF and image inputs
Output typesExtracted text, Markdown, tables, structural annotations, bounding boxes, and confidence scores
Output modalityText and structured data; it does not generate images, audio, or video
Context lengthNot presented as a conventional language-model context specification in the supplied documentation
Maximum output tokensNot specified as a conventional token limit
Tool or function useNot a tool-calling or general agent model
Batch processingSupported by the OCR API
Fine-tuningNot verified in the supplied research

The model's structured JSON capability should not be confused with general-purpose function calling. The OCR API supports JSON and structured JSON schema options for organizing OCR results, but the supplied research does not establish that OCR 4.1 independently performs arbitrary external actions or invokes tools.

Similarly, the absence of a conventional context-window or maximum-output-token figure is meaningful. OCR 4.1 is priced and documented as a page-oriented document service, not as a token-metered chat model. Applications processing large files should rely on Mistral's current API documentation and their own testing to determine practical file, page, and response constraints.

Pricing and processing economics

Mistral lists standard OCR 4.1 API pricing at $4 per 1,000 pages. Cached input is listed at $0.40 per 1,000 pages. Mistral separately lists Document AI usage at $5 per 1,000 annotated pages. These are different pricing references, so the applicable cost depends on whether an application uses standard OCR processing or Document AI annotation features.

Mistral's OCR launch materials state that batch processing receives a 50% discount. On that basis, the standard API price would be approximately $2 per 1,000 pages for eligible batch processing. Batch pricing should be confirmed against the current billing documentation before production budgeting.

Page-based pricing is easier to estimate for regular document ingestion than token pricing, especially when processing invoices, archives, forms, or large collections of PDFs. However, total project cost can also include storage, retries, validation, downstream model calls, and human review. OCR 4.1's price should therefore be evaluated as one part of a complete document pipeline.

Main strengths

  • Preserves document structure: Labels, tables, layout coordinates, and other annotations can be more useful than a plain text transcript.
  • Suitable for ingestion pipelines: Batch support and page-based pricing fit large-scale indexing and document-processing workloads.
  • Useful confidence information: Page-, block-, or word-level scores can help applications route uncertain results for review.
  • Supports visual documents: PDFs and images can be converted into data for search, retrieval, classification, and analysis.
  • Works with downstream systems: Structured output can feed enterprise search, retrieval-augmented generation, compliance tools, and workflow automation.

These strengths are most relevant when the application needs to know not only what a document says, but also how that content is organized and where it appeared.

Limitations and trade-offs

OCR 4.1 is not intended to hold open-ended conversations, write software, perform general reasoning, generate images, process audio, or understand video. A system that needs those capabilities will need another model alongside OCR 4.1 or instead of it.

The supplied model documentation does not provide a conventional context length or maximum output-token limit. That makes it inappropriate to assume that any arbitrary document can be submitted as one request without checking the current API constraints. Large-file workflows should account for page handling, response size, retries, and validation.

OCR output should also be reviewed when accuracy is consequential. Tables, unusual layouts, low-quality scans, handwriting, stamps, signatures, equations, and visually complex pages can require application-level validation or human checking. Confidence scores can help prioritize review, but they do not guarantee that every extracted field is correct.

Finally, the editorial scores associated with this entry are comparative estimates rather than Mistral-published benchmarks. The supplied research gives OCR 4.1 a high relative speed and cost score for its category, but those scores should not be read as standardized performance measurements.

Best use cases

OCR 4.1 is a strong fit for workflows that begin with visual documents and end with searchable or structured data. Practical examples include:

  • Indexing PDF archives for enterprise search.
  • Preparing document collections for retrieval-augmented generation.
  • Extracting invoice fields and preserving invoice tables.
  • Processing contracts for clauses, headings, and reference sections.
  • Digitizing forms, reports, manuals, and historical records.
  • Routing compliance documents for automated or human review.
  • Identifying signatures, equations, captions, headers, and footers.
  • Feeding document content into agents or business workflows that need typed structure.

For example, a contract pipeline could use OCR 4.1 to identify paragraphs, headings, tables, and page coordinates, then pass the resulting text and annotations to a separate language model for clause analysis. In that arrangement, OCR 4.1 performs document extraction while the other model performs interpretation.

When to choose OCR 4.1

Choose OCR 4.1 when the primary problem is converting PDFs or images into structured, searchable, machine-readable content. It is especially attractive when page-based pricing, batch processing, bounding boxes, structural labels, and confidence scores are more important than conversational interaction.

Choose a general-purpose language model instead when the input is already clean text and the main task is reasoning, drafting, coding, planning, or dialogue. Choose a vision-language model with broader reasoning capabilities when the application must interpret images in a conversational workflow rather than primarily extract document structure. Choose a specialized downstream system when a regulated workflow requires domain-specific validation that OCR alone does not provide.

For many production systems, the most appropriate design is not an either-or decision. OCR 4.1 can serve as the first stage: extract and structure the document, validate uncertain fields, and then send the resulting content to a separate model or business process for search, summarization, classification, or decision support.

Bottom line

Mistral OCR 4.1 is a page-priced OCR and document-understanding service focused on structured extraction from PDFs and images. Its distinguishing value is the combination of text extraction with layout coordinates, block types, table formatting, and confidence information. It is a practical choice for document ingestion, search, RAG preparation, forms, invoices, contracts, and archival workflows, but it should not be treated as a general-purpose chat, coding, reasoning, or agent model. Its undocumented conventional context and output-token limits also mean that implementation teams should verify current API constraints before designing large-document workloads.


Answers to Frequently Asked Questions

What are the limitations of Mistral OCR 4.1?
OCR accuracy can vary with document quality, unusual layouts, low-resolution scans, handwriting, stamps, signatures, equations, tables, and visually complex pages. Confidence scores help prioritize review but do not guarantee correctness. The supplied documentation also does not specify a conventional context window or maximum output-token limit, so large-document workflows should verify current API constraints and include validation, retries, and possible human review.
Is Mistral OCR 4.1 a general-purpose language model?
No. Mistral OCR 4.1 is a specialized OCR and document-understanding service designed to extract and structure content from documents and images. It is not intended for open-ended conversation, coding, general reasoning, image generation, audio or video processing, or arbitrary tool use. A separate language model can analyze the extracted content when needed.
How much does Mistral OCR 4.1 cost?
Standard OCR 4.1 API pricing is listed at $4 per 1,000 pages, while cached input is listed at $0.40 per 1,000 pages. Mistral separately lists Document AI usage at $5 per 1,000 annotated pages. Eligible batch processing may receive a 50% discount, reducing the standard rate to approximately $2 per 1,000 pages, subject to confirmation in the current billing documentation.
What is Mistral OCR 4.1 used for?
Mistral OCR 4.1 converts PDFs and images into searchable, machine-readable content while preserving document structure. It can extract text, tables, Markdown, bounding boxes, structural block labels, and confidence scores for workflows such as enterprise search, invoice processing, contract analysis, RAG preparation, and archival digitization.
What does Mistral OCR 4.1 return?
The service can return extracted text, Markdown, tables, paragraph-level bounding boxes, structural labels for elements such as titles, lists, equations, signatures, headers, and footers, as well as confidence scores at page, block, or word level. It also supports structured JSON output for downstream systems.


Sources 5
Provider

About Mistral AI