What is Mistral OCR 4.0?
Mistral OCR 4.0 is a specialized document-intelligence model from Mistral AI. Its canonical API model identifier is mistral-ocr-4-0. The service accepts documents and images, extracts their textual content, and returns information about the document's layout and structure.
This is an OCR model rather than a general conversational model. OCR, or optical character recognition, converts text contained in scans, photographs, PDFs, and other visual documents into machine-readable data. Mistral OCR 4.0 adds document understanding to that process: instead of returning only a continuous transcription, it can indicate what each content block represents and where it is located.
Mistral released OCR 4.0 on June 23, 2026. It is generally available through Mistral's API and Document AI offerings. OCR 4.1 is now the newer model in Mistral's OCR lineup, but OCR 4.0 remains listed as an available model and retains its documented pricing and capabilities.
Document processing capabilities
OCR 4.0 returns extracted text in reading order and can preserve useful page structure. Supported structural categories include text, titles, lists, tables, images, equations, captions, code, references, side text, headers, footers, and signatures.
For example, an invoice-processing application could distinguish the invoice title from line-item tables and totals. A legal-document pipeline could separate headings, body paragraphs, footnotes, and signatures. A search system could use block positions to show a result in its original page context instead of treating the source as an undifferentiated text file.
The model is intended for PDFs, images, presentations, word-processing documents, and OpenDocument files. Mistral's launch material describes support for approximately 170 languages across 10 language groups. This is a provider-reported coverage claim; the quality of extraction can still vary with language, scan quality, typography, layout complexity, and the source document itself.
Layout information and bounding boxes
A bounding box is a set of coordinates describing where an element appears on a page. OCR 4.0 provides paragraph-level bounding boxes and can expose structural information for individual document blocks. This allows downstream software to connect extracted text with its original visual location.
Bounding boxes are useful for citation-aware retrieval, visual document viewers, redaction, table handling, semantic chunking, and human review. A validation interface, for instance, can highlight the exact region from which an extracted invoice number was read. This is a more actionable output than plain text alone.
Structured annotations and confidence information
The OCR endpoint can return ordinary extracted text or structured annotations. Developers can request JSON objects or JSON-schema-constrained output for document-specific extraction tasks. A schema might define fields for an invoice number, supplier, date, currency, and total, allowing the returned information to fit a downstream application more predictably.
OCR 4.0 also supports confidence information at page, block, or word granularity. Confidence scores do not guarantee that an extraction is correct, but they can help applications identify uncertain regions and route them to human review. This is particularly relevant for documents where a single character error could change a payment amount, account number, legal clause, or compliance result.
Inputs, outputs, and technical limits
OCR 4.0 supports text extraction from visual and document inputs, including PDFs, images, presentations, word-processing files, and OpenDocument files. Its direct output is text and structured document data, not generated images, audio, video, or speech. The model does not provide general-purpose text generation, image generation, or conversational reasoning as its primary function.
The documented OCR response can include page-level results, markdown, extracted images, structural blocks, annotations, and usage information. The exact returned fields depend on the request and endpoint options.
Mistral does not publish a conventional context-window or maximum-output-token specification for OCR 4.0. That is consistent with its page-based document-processing design: usage and pricing are measured primarily by pages rather than by an advertised text-token context limit. Applications should therefore test the size and format of their documents against the current API behavior instead of assuming a general language-model context window.
The research does not verify streaming support, tool or function calling, or fine-tuning for OCR 4.0. It is best understood as a document extraction endpoint rather than an agent model that independently calls tools or performs multi-step reasoning.
Pricing and access
Standard API pricing is $4 per 1,000 pages. Mistral's pricing documentation also lists cached-input pricing of $0.40 per 1,000 pages. The launch announcement reports Batch API processing at $2 per 1,000 pages, which can be relevant for asynchronous, high-volume workloads.
Mistral's Document AI no-code experience is priced separately at $5 per 1,000 pages. That price should not be confused with the standard OCR 4.0 API rate. Enterprise self-hosting is also described as an option for customers with data-residency, sovereignty, or privacy requirements, but availability and commercial terms may depend on an enterprise agreement rather than the public API price.
Because the service is priced by page, cost planning differs from token-priced language models. A document with sparse text and a document packed with small print may each count as pages even though their amount of extracted text differs substantially. Teams should estimate page volume, cached usage, batch eligibility, and any separate Document AI or enterprise costs before deployment.
Strengths and limitations
Where OCR 4.0 is strong
- Structured extraction: It can preserve document blocks such as headings, lists, tables, equations, captions, and signatures rather than returning only a flat transcription.
- Spatial information: Paragraph-level bounding boxes connect extracted content to its location on the page.
- Downstream automation: JSON and JSON Schema support makes it suitable for extracting fields into business systems.
- Quality review: Page-, block-, and word-level confidence information can help prioritize manual verification.
- Multilingual document coverage: Mistral describes support for approximately 170 languages and 10 language groups.
- High-volume economics: Page-based pricing, caching, and Batch API processing provide options for large ingestion workloads.
Important limitations
- Not a general AI assistant: OCR 4.0 is not intended for open-ended conversation, coding assistance, image generation, or general reasoning.
- No published context or output-token limits: Mistral does not provide standard context-window and maximum-output-token values for this page-based service.
- Extraction still requires validation: Confidence scores can identify uncertainty but do not eliminate recognition errors.
- Limited verified model features: Tool use, streaming, and fine-tuning are not verified in the supplied specifications.
- Model succession: OCR 4.1 is the newer Mistral OCR model, so new integrations may need to compare its capabilities and migration implications with OCR 4.0.
These limitations matter most in high-stakes workflows. OCR 4.0 should not be used as the sole authority for medical diagnosis, legal conclusions, financial decisions, or safety-critical automation. Extracted information should be checked when an error could cause material harm.
Best use cases for Mistral OCR 4.0
OCR 4.0 is a good fit when an organization needs to ingest many documents and retain enough structure for later processing. Practical uses include:
- Enterprise search: Index text while retaining page locations, headings, tables, and other document structure.
- Retrieval-augmented generation: Create more meaningful chunks for a question-answering system and preserve source locations for citations.
- Invoice and form processing: Extract selected fields from semi-structured documents into accounting or workflow systems.
- Compliance review: Identify relevant sections, signatures, headers, footers, and uncertain passages for review.
- Knowledge-base construction: Convert document collections into searchable, structured records.
- Document automation: Feed page-aware OCR results into classification, validation, redaction, or approval workflows.
It is especially appropriate when the location and type of content matter as much as the words themselves. A basic text-only OCR system may be cheaper or simpler for straightforward transcription, while a general multimodal language model may be more appropriate when the task requires broad reasoning over a small number of documents. OCR 4.0's main trade-off is specialization: it offers document-oriented structure and page-based processing rather than a broad conversational feature set.
When to choose Mistral OCR 4.0
Choose OCR 4.0 when your primary problem is converting a large or varied document collection into structured, searchable data. Its bounding boxes, block labels, confidence information, and schema-constrained annotations are more directly useful than plain text extraction for enterprise workflows.
It may be a sensible option when predictable page-based pricing, Batch API processing, multilingual coverage, or an enterprise deployment path is important. The standard $4-per-1,000-page rate can also be attractive for high-volume ingestion, although actual cost depends on page count, caching, batching, and the chosen Mistral product.
Consider OCR 4.1 instead when starting a new integration and the latest model capabilities are more important than keeping an OCR 4.0 implementation. Consider a general multimodal model when the central task is reasoning, summarization, question answering, or interactive analysis rather than systematic OCR across a document collection. Consider a simpler OCR tool when you only need a basic transcription and do not need structural labels, bounding boxes, confidence granularity, or JSON extraction.
Overall assessment
Mistral OCR 4.0 is best evaluated as a structured document-ingestion component, not as a general-purpose AI model. Its distinguishing value is the combination of extracted text, document structure, page coordinates, confidence information, and structured annotations. Those outputs can reduce the work required to build search, retrieval, extraction, and review systems around PDFs and other complex documents.
The main decision is whether those capabilities justify using a specialized OCR service when a newer OCR model or a broader multimodal option may be available. For existing workflows that need OCR 4.0's documented behavior and page-based pricing, it remains a practical choice. For new deployments, teams should compare it directly with OCR 4.1 and validate representative documents before committing to a production migration or high-stakes automated process.

