PaddleOCR-VL

PaddleOCR-VL-1.5

by Baidu · Available; superseded by PaddleOCR-VL-1.6

PaddleOCR-VL-1.5 is a compact Apache 2.0 vision-language model for multilingual document parsing. It recognizes text, tables, formulas, charts, seals, layouts, and irregular regions, with improvements for skewed, warped, scanned, screen-photographed, and unevenly illuminated documents. It can be deployed locally through PaddleOCR, although PaddleOCR-VL-1.6 is now the newer compatible successor.

Text Reasoning Coding
PaddleOCR-VL-1.5 is a compact vision-language model built for document understanding rather than general-purpose conversation. Released on January 29, 2026, it combines visual document analysis with text instructions and can extract structured information from complex pages while requiring substantially fewer deployment resources than much larger general-purpose vision-language models.
Outputs

What PaddleOCR-VL-1.5 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

3/10 Reasoning
2/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family PaddleOCR-VL
Model type Multimodal
Context window 131K tokens
Release date 2026-01-29
Status Available; superseded by PaddleOCR-VL-1.6
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for this exact model in the reviewed model card, technical report, or configuration files.

Model notes

Open-weight Apache 2.0 model with approximately 0.9B parameters. The official model repository is PaddlePaddle/PaddleOCR-VL-1.5, while the inference configuration names the model PaddleOCR-VL-1.5-0.9B. PaddlePaddle released PaddleOCR-VL-1.6 as a newer compatible successor, but PaddleOCR-VL-1.5 remains downloadable and usable. Provider-reported performance includes 94.5% on OmniDocBench v1.5. The documented 131,072-token limit comes from the Hugging Face configuration's max_position_embeddings field and is not a guarantee of a fixed number of document pages.

Model guide

PaddleOCR-VL-1.5: Compact Open-Weight Model for Real-World Document Parsing

PaddleOCR-VL-1.5 is PaddlePaddle's approximately 0.9-billion-parameter, open-weight vision-language model for multilingual document parsing. It is designed to recognize and structure text, tables, formulas, charts, seals, layouts, and irregular document regions, including content found in skewed, warped, screen-photographed, scanned, or unevenly illuminated documents.

What is PaddleOCR-VL-1.5?

PaddleOCR-VL-1.5 is an open-weight vision-language model from PaddlePaddle, distributed through the PaddleOCR ecosystem and the PaddlePaddle Hugging Face organization. Its primary job is document parsing: turning document images into recognized text, detected regions, and structured information that software can further process.

Unlike a general conversational model, it is optimized for pages rather than open-ended dialogue. It is intended to identify and interpret ordinary text, tables, mathematical formulas, charts, document layouts, seals, and text regions that require both localization and recognition. The model can receive document images alongside textual instructions and produces text or structured parsing results through the PaddleOCR workflow.

The model was released on January 29, 2026. PaddleOCR-VL-1.6 is now the newer successor, but version 1.5 remains downloadable and usable. The 1.6 documentation states that its architecture remains compatible with version 1.5, which can make a later migration possible without a substantial redesign of an existing application.

Core document capabilities

PaddleOCR-VL-1.5 is designed for documents where simple text extraction is not enough. It attempts to preserve the relationships between page elements and to recognize content that traditional OCR systems may treat poorly.

  • Multilingual text recognition: extracts text from documents in multiple languages.
  • Layout analysis: identifies document structures and page elements rather than treating the entire page as an undifferentiated image.
  • Table recognition: parses tables and supports merging table content across pages.
  • Formula recognition: handles mathematical expressions as a specific document element.
  • Chart understanding: analyzes charts and their surrounding document context.
  • Text spotting: localizes text lines and recognizes their contents.
  • Seal recognition: identifies seals, an important requirement in many official and business documents.
  • Irregular-shaped localization: detects document elements that do not fit simple rectangular regions.
  • Cross-page reconstruction: helps reconnect paragraphs, headings, and tables that continue from one page to another.

These capabilities make the model more suitable for invoices, reports, forms, academic papers, administrative records, scanned contracts, and similar material than a basic OCR endpoint that only returns a sequence of characters.

Designed for difficult document images

A major focus of the 1.5 release is robustness outside controlled scanning conditions. Real documents are often photographed with a phone, displayed on a screen, captured at an angle, or affected by poor lighting. Pages may also be skewed, warped, or contain scanning artifacts.

PaddlePaddle reports that PaddleOCR-VL-1.5 improves handling of skew, warping, screen photography, scanning artifacts, and complex illumination. The model was evaluated on Real5-OmniDocBench, a benchmark intended to measure these real-world conditions. According to PaddlePaddle's release documentation, it achieved 94.5% on OmniDocBench v1.5.

The 94.5% figure is a provider-reported benchmark result, not a guarantee for every document. Results can vary with language, image quality, typography, layout complexity, handwriting, page composition, and the quality of the surrounding parsing pipeline. It is most useful as evidence of the model's intended positioning: difficult document images are a central target, not an incidental use case.

Technical specifications and deployment

PaddleOCR-VL-1.5 has approximately 0.9 billion parameters. A parameter is a learned numerical value used by the model; fewer parameters generally make a model easier to run than a much larger model, although actual resource requirements also depend on image size, precision, batch size, and the inference software.

The official Hugging Face configuration specifies a maximum position-embedding length of 131,072 tokens and bfloat16 weights. This is a model configuration limit for token positions, not a promise that every document of a particular number of pages will fit. Page count depends on image resolution, extracted visual tokens, instructions, and the output generated by the parsing pipeline.

The official model repository is PaddlePaddle/PaddleOCR-VL-1.5, while the inference configuration identifies the underlying model as PaddleOCR-VL-1.5-0.9B. The model is available under the Apache 2.0 license. The recommended route is to use the PaddleOCR document-parser pipeline, which provides the surrounding processing needed to turn model output into useful document results.

Depending on the current PaddleOCR release and deployment path, supported integration options include PaddlePaddle, Transformers-compatible loading, vLLM, SGLang, FastDeploy, MLX-VLM, and llama.cpp-related routes. Support details can differ between backends, so users should verify the current PaddleOCR documentation before selecting a production runtime.

Inputs, outputs, and supported functions

The model accepts text and document images. Its documented input profile supports text input and image input; audio and video input are not identified for this model. Its output is text or structured document-parsing information rather than directly generated images, audio, or video.

PaddleOCR-VL-1.5 should therefore be understood as a multimodal input model with text-oriented output. The visual part of the system interprets page images, while the resulting recognition and layout information is returned in a form that an application can display, store, search, or transform.

The supplied model information does not verify provider-level tool calling, function calling, streaming, fine-tuning, caching, batch API access, or a separate JSON-schema mode. Structured results can be produced through the PaddleOCR pipeline, but that should not be confused with a general-purpose provider API offering a formally documented JSON mode. Applications that need strict schemas may need to validate and post-process the parser's output.

Speed, cost, and capability trade-offs

With approximately 0.9 billion parameters, PaddleOCR-VL-1.5 is positioned as a relatively compact document model. Its smaller size can reduce deployment cost and improve throughput compared with using a large general-purpose vision-language model for every page. The model records supplied with this review rate its speed and cost favorably as editorial evaluations, not as provider-published guarantees.

The trade-off is specialization. A compact document parser may be a better fit for high-volume OCR and layout extraction than a large conversational model, but it is not intended to replace a general model for broad reasoning, open-ended dialogue, software development, web search, speech, or media generation. Its reasoning capability is primarily applied to interpreting document structure and relationships between page elements. Its coding capability is limited because code generation is not its purpose.

There is no verified hosted price for PaddleOCR-VL-1.5 in the supplied information. Because it is open-weight and locally deployable, costs depend on hardware, inference infrastructure, image volume, and the selected serving backend rather than on a documented per-token price for this exact model.

When to choose PaddleOCR-VL-1.5

Choose PaddleOCR-VL-1.5 when the main problem is extracting information from visually complex documents and you want an open-weight model that can be deployed through PaddleOCR. It is particularly relevant when documents include tables, formulas, charts, seals, multiple languages, continuing sections across pages, or images captured under imperfect conditions.

  • Choose it for: multilingual OCR, document layout analysis, table extraction, formula recognition, chart understanding, seal detection, scanned pages, warped pages, and screen photographs.
  • Choose it when: local deployment, Apache 2.0 licensing, and control over the inference environment are important.
  • Choose a larger general-purpose vision-language model instead when: you need broad conversation, extensive general reasoning, coding assistance, web search, or tool use beyond document parsing.
  • Choose a hosted OCR service instead when: you prefer a managed endpoint and its pricing, data-handling terms, and operational simplicity are more important than local model control.
  • Consider PaddleOCR-VL-1.6 when: you want the newer successor and your deployment can move to the compatible current architecture.

For production use, test representative pages from the actual document collection rather than relying only on benchmark scores. Include low-quality scans, unusual fonts, tables spanning pages, pages with seals, and the languages that matter to the application. Also validate the output format, because accurate recognition does not automatically guarantee a clean database-ready schema.

Limitations and current status

PaddleOCR-VL-1.5 is specialized rather than universal. It does not provide verified support for audio or video understanding, image generation, video generation, speech, web search, or general tool execution. The available information also does not establish a fixed maximum output-token value, a hosted API price, or a provider-level JSON-schema feature.

The 131,072-token configuration should not be interpreted as a fixed page limit, and the reported 94.5% benchmark result should not be treated as a universal accuracy rate. Performance depends on the input and the complete OCR and parsing pipeline.

In the current PaddleOCR lineup, version 1.5 is best viewed as an available, compact predecessor to PaddleOCR-VL-1.6. It remains useful where its open-weight license, deployment compatibility, and document-focused capabilities meet project requirements, but new projects should compare it with the newer version before standardizing on it.


Answers to Frequently Asked Questions

What inputs and outputs does PaddleOCR-VL-1.5 support?
The model accepts text and document images and produces text or structured document-parsing information. It is intended for visual document understanding rather than audio, video, image generation, speech, or general-purpose conversational tasks.
Should I choose PaddleOCR-VL-1.5 or PaddleOCR-VL-1.6?
PaddleOCR-VL-1.6 is the newer successor and should be compared first for new projects. PaddleOCR-VL-1.5 remains useful when its compact size, Apache 2.0 license, local deployment options, or existing integration requirements are important. The 1.6 documentation states that its architecture remains compatible with version 1.5.
What are the technical specifications and license of PaddleOCR-VL-1.5?
PaddleOCR-VL-1.5 has approximately 0.9 billion parameters, uses bfloat16 weights, and specifies a maximum position-embedding length of 131,072 tokens. Its official repository is PaddlePaddle/PaddleOCR-VL-1.5, and it is available under the Apache 2.0 license.
What is PaddleOCR-VL-1.5 used for?
PaddleOCR-VL-1.5 is an open-weight vision-language model designed for document parsing. It extracts text, analyzes layouts, recognizes tables and mathematical formulas, understands charts, detects seals, and reconstructs content across pages.
How well does PaddleOCR-VL-1.5 handle low-quality or difficult document images?
The model is designed to handle skewed or warped pages, screen photographs, scanning artifacts, poor lighting, and other real-world image problems. PaddlePaddle reports a 94.5% result on OmniDocBench v1.5, although actual accuracy depends on the documents and the complete parsing pipeline.


Sources 5
Provider

About Baidu