Falcon Perception

Falcon OCR

by Technology Innovation Institute (TII) · Current open-weight release

Falcon OCR is a compact open-weight vision-language model for document image extraction. It produces plain text, LaTeX formulas, and HTML tables, supports optional layout detection, and can run through Transformers, PyTorch, MLX, vLLM, or Docker. Its small size suits local and cost-conscious deployments, while degraded scans, tiny text, limited general reasoning, and the absence of official hosted pricing remain important considerations.

Text Reasoning Coding
Falcon OCR is a specialized document-understanding model rather than a general-purpose chatbot. Given a document image, it can extract ordinary text, recognize mathematical formulas as LaTeX, and convert tables into HTML. Its relatively small size, Apache-2.0 license, and support for local deployment make it relevant for developers and organizations that need OCR without relying on a hosted commercial API.
Outputs

What Falcon OCR can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming Batch API
Model profile

Performance characteristics

3/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Falcon Perception
Model type Multimodal
Context window 16K tokens
Release date 2026-03-28
Status Current open-weight release
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was published in the model card, configuration, repository documentation, or cited technical report.

Model notes

Falcon OCR is a 300M-parameter early-fusion vision-language model. The canonical Hugging Face identifier is tiiuae/Falcon-OCR. It accepts document images and can return plain text, LaTeX for formulas, or HTML for tables. The model card lists category-specific modes including text, table, formula, caption, footnote, list-item, page-footer, page-header, section-header, and title. An optional two-stage layout pipeline uses PP-DocLayoutV3 for region detection. The repository provides Transformers usage, a PyTorch inference engine, Apple Silicon MLX support, a vLLM-compatible server, and a Docker deployment. The model is open-weight under the Apache-2.0 license; no official hosted token pricing was identified. The model card reports 80.3% on olmOCR-Bench and 88.64 on OmniDocBench, while noting weaknesses on degraded scans and very small text. The repository documentation indicates that Transformers usage requires custom code and PyTorch 2.5 or newer for FlexAttention. Falcon OCR is distinct from the existing Falcon Perception model.

Model guide

Falcon OCR: An Open-Weight Model for Efficient Document Text, Formula, and Table Extraction

Falcon OCR is a compact, open-weight 300-million-parameter vision-language model from the Technology Innovation Institute, designed specifically for document OCR. It converts document images into plain text, LaTeX formulas, or HTML tables, with optional layout-aware parsing for structured documents.

What is Falcon OCR?

Falcon OCR is a 300-million-parameter vision-language model developed by the Technology Innovation Institute (TII). It is intended for optical character recognition (OCR): reading text and structured content from document images. The model is available as open weight through the tiiuae/Falcon-OCR repository on Hugging Face and is licensed under Apache-2.0 according to the supplied model documentation.

Unlike a general conversational model, Falcon OCR is focused on turning visual documents into usable text representations. Depending on the selected mode, an image can produce ordinary text, LaTeX for mathematical formulas, or HTML for tables. The model documentation also identifies specialized categories such as captions, footnotes, list items, page headers, page footers, section headers, and titles.

This makes Falcon OCR better understood as a compact document-processing component than as a standalone assistant. It can be integrated into an application, run locally, or served through deployment tools documented in its repository.

How the model processes documents

Falcon OCR uses an early-fusion vision-language architecture. Image patches and text tokens are passed through a shared Transformer from the first layer, rather than being processed by entirely separate vision and language stacks that are combined only later. In practical terms, this allows the model to connect visual regions with the text representation it generates.

The core model is intended for document OCR, while an optional two-stage layout pipeline can add region detection before recognition. The supplied documentation identifies PP-DocLayoutV3 for this purpose. Layout detection is useful when a page contains multiple logical areas, such as a title, paragraphs, a table, a formula, and page-footer text. Instead of treating the entire page as one undifferentiated image, the pipeline can identify regions and apply the appropriate extraction mode.

The available repository materials describe several deployment routes: Transformers-based inference, a PyTorch inference engine, Apple Silicon MLX support, a vLLM-compatible server, and Docker deployment. Transformers use requires custom code, and the documentation notes a requirement for PyTorch 2.5 or newer when FlexAttention is used.

Supported inputs and outputs

Falcon OCR accepts document images and produces text-based results. It does not generate images, audio, or video. Its multimodal capability is therefore visual input combined with textual output.

CapabilitySupported behavior
Primary inputDocument images
Text extractionPlain text
Formula recognitionLaTeX output
Table extractionHTML table output
Layout processingOptional region detection and layout-aware parsing
Audio, video, and image generationNot supported

Output format matters here. Plain text is suitable for searchable archives, notes, and document indexing. LaTeX is more useful when mathematical notation must remain editable or renderable. HTML tables can preserve rows and columns more effectively than a plain-text transcription, although the accuracy of any table conversion still depends on image quality and layout complexity.

What Falcon OCR is useful for

The model is a good fit for document pipelines that need more than simple character recognition. Suitable examples include receipts, invoices, academic papers, forms, scanned reports, and pages containing mathematical notation or tables.

  • Searchable document archives: Convert scanned pages into text that can be indexed and searched.
  • Invoice and receipt processing: Extract printed content from financial documents for later application-level parsing.
  • Academic and technical papers: Recover body text, formulas, headings, footnotes, and tables from page images.
  • Structured table extraction: Produce HTML tables for downstream cleaning or conversion into application data.
  • Private or offline workflows: Run the model on controlled infrastructure when documents should not be sent to a third-party OCR service.
  • Edge-oriented deployments: Use its comparatively compact size where a smaller model is more practical than a large general vision-language system.

These are use-case recommendations based on the model's documented functions, not a guarantee that every document type will be extracted perfectly. Production systems should validate important fields, particularly when OCR results feed accounting, compliance, or other high-consequence processes.

Accuracy and practical limitations

The model card reports 80.3% on olmOCR-Bench and 88.64 on OmniDocBench. These are provider or model-card-reported benchmark figures; they should not be treated as universal accuracy guarantees because benchmark composition, formatting, image quality, and evaluation methods affect results.

The same documentation notes weaknesses with degraded scans and very small text. This is an important limitation for users working with old photocopies, low-resolution archives, compressed images, or dense pages. A layout-aware pipeline may help separate regions, but it cannot fully recover information that is missing or illegible in the source image.

Falcon OCR is also not positioned as a general-purpose reasoning or coding model. Its documented purpose is extraction and document parsing. It may return text from an image, but that does not make it a replacement for a model designed for extended reasoning, software development, open-ended conversation, or broad multimodal analysis. It also has no documented native tool or function-calling support.

Context, speed, and cost

The supplied model record lists a context length of 16,384 tokens. This is the documented context figure, but no separate maximum output-token limit is identified. Users building a pipeline should therefore test page size, prompt format, and expected output length rather than assuming that a large multi-page document will always fit in one request.

Falcon OCR has no official hosted token price identified in the supplied research. The practical cost is therefore determined by the infrastructure used to run it: local hardware, a private server, a vLLM deployment, Docker infrastructure, or a third-party hosting service. The open-weight Apache-2.0 distribution can reduce software licensing costs, but it does not make compute, storage, engineering, or operational costs disappear.

Its 300-million-parameter size is a meaningful trade-off. A smaller specialized model can be faster and less expensive to operate than a large general-purpose vision-language model, especially for high-volume OCR. In exchange, it is narrower in scope and may be less capable on difficult visual reasoning, ambiguous layouts, or tasks that require extended interpretation beyond transcription. The supplied editorial assessment rates its speed and cost favorably, but those ratings are evaluations rather than provider-published benchmark categories.

Reasoning, coding, and tool support

Falcon OCR has limited reasoning relevance. It can interpret document regions sufficiently to classify or extract their contents, but the research does not describe it as a reasoning-first model. Its primary value is faithful conversion of visual document content into structured text formats.

Coding capability is not a target feature. The model can output HTML tables and LaTeX, which are structured or markup-based representations, but that should not be confused with general code generation. There is also no documented built-in web search, function calling, action execution, or external tool-use capability.

Developers can still place Falcon OCR inside a larger workflow. For example, an application could send an image to the model, validate the returned HTML or text, and then pass the result to a separate parser, database, or business process. Those surrounding capabilities would come from the application, not from Falcon OCR itself.

When to choose Falcon OCR

Choose Falcon OCR when the central problem is document image extraction and you value local control, open weights, and a relatively compact model. It is especially appropriate when you need one or more of the following:

  • Local or self-hosted OCR instead of a hosted commercial endpoint.
  • Plain text, LaTeX formulas, or HTML tables from document images.
  • Integration with a Python, PyTorch, Transformers, MLX, vLLM, or Docker-based workflow.
  • A smaller model for high-volume processing or constrained infrastructure.
  • Apache-2.0 licensing and the ability to inspect or integrate the model weights.

Another option may be more appropriate if you need a polished managed API, guaranteed service-level availability, broad document-management features, or extensive support for difficult multimodal reasoning. A general vision-language model may be preferable for tasks that combine OCR with summarization, visual question answering, complex interpretation, or conversation. A conventional OCR service may be preferable when turnkey scaling, vendor-managed operations, and standardized enterprise support matter more than self-hosting.

Availability and final assessment

Falcon OCR is available as an open-weight model through TII's Falcon repository on Hugging Face. The supplied research identifies it as a current open-weight release and places it within the Falcon Perception model family, while also distinguishing it from the broader Falcon Perception model itself.

Its main distinction is specialization. Falcon OCR does not try to be an all-purpose assistant; it concentrates on converting document images into useful textual and semi-structured formats. That focus, combined with its compact parameter count and multiple deployment paths, makes it a credible option for developers building private or cost-conscious OCR systems. The main trade-off is that users must manage deployment and quality validation themselves, and degraded scans, tiny text, complex pages, and general-purpose reasoning remain important limitations.


Answers to Frequently Asked Questions

When should developers choose Falcon OCR?
Developers should consider Falcon OCR when they need local or self-hosted document OCR, open weights under the Apache-2.0 license, plain text, LaTeX, or HTML table extraction, and a relatively compact model for cost-conscious or high-volume processing. A managed OCR API or general vision-language model may be more suitable when turnkey scaling, enterprise support, or broader multimodal reasoning is required.
What are Falcon OCR's main limitations?
Falcon OCR may perform poorly on degraded scans, very small text, low-resolution images, and complex layouts. Its reported benchmark scores are not universal accuracy guarantees, and important outputs should be validated. It is also specialized for document extraction rather than general reasoning, coding, conversation, web search, or tool use.
Does Falcon OCR support formulas and tables?
Yes. Falcon OCR can output mathematical formulas as LaTeX and tables as HTML. An optional layout pipeline using PP-DocLayoutV3 can detect regions such as tables, formulas, titles, paragraphs, headers, and footers before applying the appropriate extraction mode.
What is Falcon OCR and what can it extract from document images?
Falcon OCR is a 300-million-parameter open-weight vision-language model from the Technology Innovation Institute (TII). It extracts document content as plain text, mathematical formulas in LaTeX, and tables in HTML, with support for elements such as headings, captions, footnotes, and lists.
What deployment options are available for Falcon OCR?
Falcon OCR can be deployed through Transformers, a PyTorch inference engine, Apple Silicon MLX, a vLLM-compatible server, or Docker. Its Hugging Face repository is tiiuae/Falcon-OCR, and the model documentation notes that Transformers use requires custom code and may require PyTorch 2.5 or newer when FlexAttention is enabled.


Sources 4
Provider

About Technology Innovation Institute (TII)