GLM-OCR

GLM-OCR

by Z.ai · Current; open-weight; available through hosted API and self-hosted deployment

GLM-OCR is a 0.9-billion-parameter multimodal OCR model from Z.ai for processing images and PDFs. It focuses on text, tables, formulas, handwriting, layout-aware parsing, and structured information extraction, with a hosted API priced at $0.03 per million input and output tokens and local deployment options.

Text Reasoning Coding
GLM-OCR is a specialized document-understanding model from Z.ai and Zhipu AI. Rather than serving as a general chatbot, it focuses on turning scanned documents, photographs, screenshots, and PDFs into searchable text, Markdown, tables, formulas, or structured JSON. Its relatively small 0.9-billion-parameter design is intended to reduce inference cost and latency, while its hosted API and open-weight release give teams a choice between managed processing and local deployment.
Outputs

What GLM-OCR can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning Structured output
Model profile

Performance characteristics

2/10 Reasoning
2/10 Coding
9/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family GLM-OCR
Model type Other
Release date 2026-02-03
Status Current; open-weight; available through hosted API and self-hosted deployment
Knowledge cutoff notes

No authoritative knowledge-cutoff date was identified for this specialized OCR model. Its purpose is document recognition rather than broad factual question answering.

Model notes

GLM-OCR is a specialized 0.9B-parameter OCR model rather than a general-purpose language model. The official hosted API supports PDF and image inputs, with a documented maximum of 10 MB per image, 50 MB per PDF, and 100 pages per document. The API returns text, Markdown, image links, HTML-formatted table content, and structured JSON depending on the task. The model is released under the MIT License. The complete official OCR pipeline also uses PP-DocLayout-V3, which has separate licensing and deployment considerations. Editorial scores are comparative estimates for OCR-focused use rather than general LLM capability.

Cost

Model pricing

Input $0.03 per 1 million tokens
Output $0.03 per 1 million tokens
Model guide

GLM-OCR: Z.ai’s Compact Model for Accurate Document Parsing

GLM-OCR is Z.ai’s 0.9-billion-parameter open-weight multimodal OCR model for extracting text, formulas, tables, handwriting, and structured information from images and PDFs. It combines a compact GLM-V architecture with layout analysis and parallel recognition to provide an efficient alternative to larger vision-language models for document-processing workloads.

What is GLM-OCR?

GLM-OCR is a multimodal optical character recognition model developed by Z.ai and Zhipu AI. OCR normally means converting text in an image or document into machine-readable characters, but GLM-OCR is designed for a broader document-understanding workflow. It can recognize ordinary text as well as mathematical formulas, tables, forms, receipts, certificates, seals, code-heavy documents, and handwritten content.

The model is positioned as a focused member of Z.ai’s GLM lineup, not as a general-purpose conversational or reasoning model. Its primary job is to interpret visual documents and return usable digital content. Depending on the task and API configuration, that content can be plain text, Markdown, image links, HTML-formatted table content, or structured JSON for key-value extraction.

At approximately 0.9 billion parameters, GLM-OCR is considerably smaller than many general-purpose vision-language models. That compact scale is important for high-volume OCR, lower-resource deployments, and applications where document throughput matters more than broad conversational ability.

How GLM-OCR processes documents

GLM-OCR combines several components rather than relying only on a language decoder. Its reported architecture includes a CogViT visual encoder, a lightweight cross-modal connector, and a GLM-0.5B language decoder. The visual encoder examines the document image, the connector passes visual information into the language component, and the decoder produces the requested textual or structured result.

The broader official OCR pipeline also uses PP-DocLayout-V3 for layout analysis. Layout analysis identifies regions such as headings, paragraphs, tables, figures, formulas, and form fields before recognition is applied to those regions. The pipeline then performs parallel recognition, which helps explain why GLM-OCR is aimed at structured documents rather than only at isolated lines of text.

This distinction matters for practical deployments. The base recognition model and the complete document-processing pipeline are not necessarily the same thing. Teams using the open-weight model locally may need to configure layout detection, region handling, preprocessing, batching, and output formatting themselves. The full workflow therefore involves more than downloading model weights and sending an image to a decoder.

Accuracy and throughput

Z.ai reports a score of 94.62 on OmniDocBench V1.5 and states that GLM-OCR ranked first overall at launch. The technical material also reports strong results for text recognition, mathematical formula recognition, table parsing, and key-information extraction. These are provider or project claims and should be interpreted in the context of the benchmark’s documents, metrics, and test conditions rather than as a guarantee for every document type.

Z.ai also reports approximately 1.86 pages per second for PDF documents and 0.67 images per second under its stated test conditions. Actual performance will depend on image resolution, document complexity, hardware, batching, preprocessing, the surrounding layout pipeline, and whether the hosted API or a local deployment is used.

The practical trade-off is straightforward: GLM-OCR gives up general-purpose language and reasoning breadth in exchange for a smaller model focused on document recognition. For a workload dominated by invoices, scanned forms, tables, or PDFs, that specialization can be more useful than paying for a much larger model with capabilities the application does not need.

Supported inputs and output formats

The hosted GLM-OCR service accepts PDF files and common image formats including JPG and PNG. Z.ai documents a maximum single-image size of 10 MB, a maximum PDF size of 50 MB, and support for documents up to 100 pages. These limits apply to the documented hosted workflow and should be checked before submitting large or unusually complex files.

The model supports Chinese, English, French, Spanish, Russian, German, Japanese, Korean, and other languages according to the supplied documentation. Recognition quality can vary with language, font, scan quality, handwriting, page layout, and the amount of visual noise in the source document.

GLM-OCR can return several forms of output:

  • Plain text for basic transcription and search indexing.
  • Markdown for preserving a readable document structure.
  • HTML-formatted table content when table layout needs to be retained.
  • Image links for document regions or related visual content.
  • Structured JSON for extracting fields such as invoice values, names, dates, identifiers, or other key information.

Its output is text or structured data rather than newly generated images, audio, or video. The model record identifies text output and structured output, but does not establish a separate general-purpose JSON mode with guarantees equivalent to a broad language-model structured-output system.

Pricing and access

The documented hosted API price is $0.03 per million input tokens and $0.03 per million output tokens. This is the listed GLM-OCR API rate, not a subscription price for the wider Z.ai consumer service. The practical cost of a document-processing job will depend on how the service tokenizes the submitted content and generated result.

GLM-OCR is also released as an open-weight model under the MIT License. Z.ai identifies Hugging Face and ModelScope as distribution channels, and the model can be deployed with tools and frameworks including vLLM, SGLang, Ollama, and Transformers. Self-hosting can provide greater control over data handling, concurrency, and infrastructure costs, but it also transfers responsibility for compatible hardware, installation, scaling, monitoring, and the supporting layout-analysis pipeline.

The complete official OCR workflow includes PP-DocLayout-V3. Its licensing and deployment terms should be reviewed separately from the MIT license associated with the GLM-OCR model weights. This is especially important when a production system uses the full pipeline rather than the recognizer alone.

Capabilities and boundaries

GLM-OCR is strongest when the input is a visual document and the desired result is a faithful transcription or structured extraction. It can help process scanned reports, receipts, invoices, certificates, forms, identity documents, screenshots, academic pages, technical material, and handwritten notes. It is also suitable for converting documents into normalized content for search, databases, or retrieval-augmented generation systems.

It is not designed to replace a general-purpose assistant. The supplied research does not describe GLM-OCR as a model for open-ended conversation, advanced reasoning, software development, web search, audio processing, video processing, or image generation. Its reasoning and coding scores are low comparative editorial ratings, not provider-published benchmark results. In practical terms, the model may extract code from a document, but that does not make it a coding assistant; it may recognize a mathematical formula, but that does not establish advanced mathematical reasoning.

The model record lists image input and text output, with no audio or video input and no direct image, audio, or video output. It also lists tool use as unsupported. There is no supplied verified context-window value, maximum output-token value, streaming specification, caching feature, or batch API guarantee. Applications should therefore avoid assuming that GLM-OCR has the same API behavior as a general-purpose language model.

Main strengths and limitations

Where GLM-OCR is strong

  • Document specialization: The model targets text, tables, formulas, handwriting, layouts, and key information rather than treating every image as a generic visual question.
  • Compact deployment: Its approximately 0.9B parameter scale is suitable for teams seeking a smaller local model or a lower-cost high-volume OCR component.
  • Structured results: Markdown, HTML tables, and JSON-style extraction can reduce the work required after recognition.
  • Open-weight flexibility: The MIT-licensed model can be downloaded and deployed locally using several supported frameworks.
  • Managed and self-hosted options: Teams can use the hosted API for simpler integration or operate their own pipeline when data control and infrastructure customization are priorities.

Limitations to plan for

  • Not a general assistant: It is not intended for broad reasoning, coding, conversation, web search, or media generation.
  • Pipeline complexity: High-quality document parsing may require layout detection and other components in addition to the base model.
  • Document limits: The hosted documentation specifies 10 MB per image, 50 MB per PDF, and up to 100 pages per document.
  • Unknown general LLM limits: A context length and maximum output-token limit were not identified in the supplied research.
  • Variable real-world accuracy: Benchmark results and reported throughput do not guarantee the same performance on poor scans, unusual layouts, dense handwriting, or specialized forms.
  • Operational responsibility when self-hosting: Local deployment requires suitable hardware and management of the layout-analysis and inference stack.

When to choose GLM-OCR

Choose GLM-OCR when the central problem is converting visual documents into reliable text or structured records. It is a particularly plausible choice for invoice and receipt extraction, searchable archives, table recovery, formula transcription, form processing, document ingestion, and OCR services that need to handle many files without deploying a very large vision-language model.

The hosted API is the simpler option when a team wants to avoid operating model infrastructure. The open-weight release is more attractive when documents must remain within a controlled environment, when concurrency needs to be tuned locally, or when the team wants to integrate recognition with its own preprocessing and layout pipeline.

Another type of vision-language model may be more appropriate when the application requires broad visual question answering, multi-step reasoning about an image, conversational interaction, coding assistance, web-connected workflows, or image generation. A larger general-purpose model may also be preferable when OCR is only one small part of a wider assistant. Conversely, a simpler traditional OCR engine may be sufficient for clean, single-column printed text and could be easier to operate if tables, formulas, handwriting, and complex layouts are not important.

Bottom line

GLM-OCR is best understood as a focused document-recognition component rather than a small general-purpose chatbot. Its combination of a compact architecture, layout-aware processing, structured output formats, reported benchmark performance, hosted API, and open-weight deployment makes it relevant for document-heavy applications. The main decision is whether its specialized OCR capability matches the workload: for extracting and organizing information from images and PDFs, it offers a clear speed-and-cost-oriented alternative to larger multimodal models; for general reasoning or interactive assistance, a different model type will be a better fit.


Answers to Frequently Asked Questions

When should you choose GLM-OCR instead of a general-purpose vision-language model?
Choose GLM-OCR when the main task is extracting reliable text or structured records from documents such as invoices, receipts, forms, tables, scanned reports, or PDFs. A general-purpose vision-language model is more suitable for broad visual question answering, multi-step reasoning, coding assistance, web-connected workflows, conversation, or image generation.
Can GLM-OCR be self-hosted?
Yes. GLM-OCR is available as an open-weight model under the MIT License through platforms such as Hugging Face and ModelScope. It can be deployed with vLLM, SGLang, Ollama, or Transformers, although production use may also require a layout-analysis component such as PP-DocLayout-V3.
How accurate and fast is GLM-OCR?
Z.ai reports a score of 94.62 on OmniDocBench V1.5 and approximately 1.86 PDF pages per second or 0.67 images per second under stated test conditions. Actual accuracy and speed vary with document complexity, image quality, hardware, preprocessing, batching, and deployment method.
What is GLM-OCR designed to do?
GLM-OCR is a compact multimodal document-recognition model from Z.ai and Zhipu AI. It converts PDFs and images into text or structured data while handling tables, formulas, forms, receipts, certificates, seals, code-heavy documents, and handwriting.
What input and output formats does GLM-OCR support?
The hosted service accepts PDF, JPG, and PNG files. It can return plain text, Markdown, HTML-formatted tables, image links, or structured JSON for extracting fields such as names, dates, invoice values, and identifiers.


Sources 5
Provider

About Z.ai