What is Mistral OCR 4.1?
Mistral OCR 4.1 is Mistral AI's current specialized optical character recognition (OCR) and document-understanding service. OCR means converting text contained in documents or images into data that software can search, index, analyze, or pass to another system. OCR 4.1 goes beyond returning one plain text stream: it is designed to preserve important parts of a document's structure.
The model accepts document and image inputs and can return extracted text, Markdown, tables, paragraph-level bounding boxes, structural block labels, and confidence scores. A bounding box identifies where an extracted element appears on the page. Structural labels indicate what a block represents, such as a title, paragraph, table, list, equation, image, caption, code block, reference, header, footer, or signature.
Its canonical API identifier is mistral-ocr-4-1. According to Mistral's changelog, mistral-ocr-latest and mistral-ocr-4 point to OCR 4.1. Mistral released the model on July 16, 2026 and marked it Generally Available on August 31, 2026. It is listed as the latest OCR model in Mistral's current catalog.
What information does it extract?
OCR 4.1 is useful when the location and role of text matter as much as the words themselves. For example, a document ingestion system can treat a heading differently from a footer, preserve a table as a table, or identify a signature block for later review.
- Text: Extracted content from supported PDFs and images.
- Layout: Paragraph-level bounding boxes and page-level organization.
- Structural blocks: Labels for text, titles, lists, tables, images, equations, captions, code, references, headers, footers, and signatures.
- Tables: Table extraction with supported Markdown or HTML formatting options.
- Confidence information: Confidence scores can be returned at page, block, or word granularity.
- Structured annotations: Machine-readable output for downstream document-processing systems.
These features make OCR 4.1 more suitable for structured extraction than a basic scanner-style OCR tool that only produces unformatted text. The exact usefulness of the output depends on the document quality, layout, language, and the way the consuming application validates the extracted data.
Where OCR 4.1 fits in Mistral AI's lineup
OCR 4.1 is a specialized model in Mistral AI's Document AI and API ecosystem. It is not a general-purpose conversational model, coding model, or autonomous agent. Its purpose is to transform visual documents into text and structured document data that can then be searched, stored, reviewed, or supplied to another model.
OCR 4.0 and OCR 3 remain separate catalog entries, but OCR 4.1 is the current canonical version for new integrations according to the supplied Mistral documentation. The service is therefore best understood as a document-ingestion component rather than a replacement for a general Mistral language model.
Inputs, outputs, and technical capabilities
| Area | Verified information |
|---|---|
| Input types | Documents and images, including PDF and image inputs |
| Output types | Extracted text, Markdown, tables, structural annotations, bounding boxes, and confidence scores |
| Output modality | Text and structured data; it does not generate images, audio, or video |
| Context length | Not presented as a conventional language-model context specification in the supplied documentation |
| Maximum output tokens | Not specified as a conventional token limit |
| Tool or function use | Not a tool-calling or general agent model |
| Batch processing | Supported by the OCR API |
| Fine-tuning | Not verified in the supplied research |
The model's structured JSON capability should not be confused with general-purpose function calling. The OCR API supports JSON and structured JSON schema options for organizing OCR results, but the supplied research does not establish that OCR 4.1 independently performs arbitrary external actions or invokes tools.
Similarly, the absence of a conventional context-window or maximum-output-token figure is meaningful. OCR 4.1 is priced and documented as a page-oriented document service, not as a token-metered chat model. Applications processing large files should rely on Mistral's current API documentation and their own testing to determine practical file, page, and response constraints.
Pricing and processing economics
Mistral lists standard OCR 4.1 API pricing at $4 per 1,000 pages. Cached input is listed at $0.40 per 1,000 pages. Mistral separately lists Document AI usage at $5 per 1,000 annotated pages. These are different pricing references, so the applicable cost depends on whether an application uses standard OCR processing or Document AI annotation features.
Mistral's OCR launch materials state that batch processing receives a 50% discount. On that basis, the standard API price would be approximately $2 per 1,000 pages for eligible batch processing. Batch pricing should be confirmed against the current billing documentation before production budgeting.
Page-based pricing is easier to estimate for regular document ingestion than token pricing, especially when processing invoices, archives, forms, or large collections of PDFs. However, total project cost can also include storage, retries, validation, downstream model calls, and human review. OCR 4.1's price should therefore be evaluated as one part of a complete document pipeline.
Main strengths
- Preserves document structure: Labels, tables, layout coordinates, and other annotations can be more useful than a plain text transcript.
- Suitable for ingestion pipelines: Batch support and page-based pricing fit large-scale indexing and document-processing workloads.
- Useful confidence information: Page-, block-, or word-level scores can help applications route uncertain results for review.
- Supports visual documents: PDFs and images can be converted into data for search, retrieval, classification, and analysis.
- Works with downstream systems: Structured output can feed enterprise search, retrieval-augmented generation, compliance tools, and workflow automation.
These strengths are most relevant when the application needs to know not only what a document says, but also how that content is organized and where it appeared.
Limitations and trade-offs
OCR 4.1 is not intended to hold open-ended conversations, write software, perform general reasoning, generate images, process audio, or understand video. A system that needs those capabilities will need another model alongside OCR 4.1 or instead of it.
The supplied model documentation does not provide a conventional context length or maximum output-token limit. That makes it inappropriate to assume that any arbitrary document can be submitted as one request without checking the current API constraints. Large-file workflows should account for page handling, response size, retries, and validation.
OCR output should also be reviewed when accuracy is consequential. Tables, unusual layouts, low-quality scans, handwriting, stamps, signatures, equations, and visually complex pages can require application-level validation or human checking. Confidence scores can help prioritize review, but they do not guarantee that every extracted field is correct.
Finally, the editorial scores associated with this entry are comparative estimates rather than Mistral-published benchmarks. The supplied research gives OCR 4.1 a high relative speed and cost score for its category, but those scores should not be read as standardized performance measurements.
Best use cases
OCR 4.1 is a strong fit for workflows that begin with visual documents and end with searchable or structured data. Practical examples include:
- Indexing PDF archives for enterprise search.
- Preparing document collections for retrieval-augmented generation.
- Extracting invoice fields and preserving invoice tables.
- Processing contracts for clauses, headings, and reference sections.
- Digitizing forms, reports, manuals, and historical records.
- Routing compliance documents for automated or human review.
- Identifying signatures, equations, captions, headers, and footers.
- Feeding document content into agents or business workflows that need typed structure.
For example, a contract pipeline could use OCR 4.1 to identify paragraphs, headings, tables, and page coordinates, then pass the resulting text and annotations to a separate language model for clause analysis. In that arrangement, OCR 4.1 performs document extraction while the other model performs interpretation.
When to choose OCR 4.1
Choose OCR 4.1 when the primary problem is converting PDFs or images into structured, searchable, machine-readable content. It is especially attractive when page-based pricing, batch processing, bounding boxes, structural labels, and confidence scores are more important than conversational interaction.
Choose a general-purpose language model instead when the input is already clean text and the main task is reasoning, drafting, coding, planning, or dialogue. Choose a vision-language model with broader reasoning capabilities when the application must interpret images in a conversational workflow rather than primarily extract document structure. Choose a specialized downstream system when a regulated workflow requires domain-specific validation that OCR alone does not provide.
For many production systems, the most appropriate design is not an either-or decision. OCR 4.1 can serve as the first stage: extract and structure the document, validate uncertain fields, and then send the resulting content to a separate model or business process for search, summarization, classification, or decision support.
Bottom line
Mistral OCR 4.1 is a page-priced OCR and document-understanding service focused on structured extraction from PDFs and images. Its distinguishing value is the combination of text extraction with layout coordinates, block types, table formatting, and confidence information. It is a practical choice for document ingestion, search, RAG preparation, forms, invoices, contracts, and archival workflows, but it should not be treated as a general-purpose chat, coding, reasoning, or agent model. Its undocumented conventional context and output-token limits also mean that implementation teams should verify current API constraints before designing large-document workloads.

