What is GLM-OCR?
GLM-OCR is a multimodal optical character recognition model developed by Z.ai and Zhipu AI. OCR normally means converting text in an image or document into machine-readable characters, but GLM-OCR is designed for a broader document-understanding workflow. It can recognize ordinary text as well as mathematical formulas, tables, forms, receipts, certificates, seals, code-heavy documents, and handwritten content.
The model is positioned as a focused member of Z.ai’s GLM lineup, not as a general-purpose conversational or reasoning model. Its primary job is to interpret visual documents and return usable digital content. Depending on the task and API configuration, that content can be plain text, Markdown, image links, HTML-formatted table content, or structured JSON for key-value extraction.
At approximately 0.9 billion parameters, GLM-OCR is considerably smaller than many general-purpose vision-language models. That compact scale is important for high-volume OCR, lower-resource deployments, and applications where document throughput matters more than broad conversational ability.
How GLM-OCR processes documents
GLM-OCR combines several components rather than relying only on a language decoder. Its reported architecture includes a CogViT visual encoder, a lightweight cross-modal connector, and a GLM-0.5B language decoder. The visual encoder examines the document image, the connector passes visual information into the language component, and the decoder produces the requested textual or structured result.
The broader official OCR pipeline also uses PP-DocLayout-V3 for layout analysis. Layout analysis identifies regions such as headings, paragraphs, tables, figures, formulas, and form fields before recognition is applied to those regions. The pipeline then performs parallel recognition, which helps explain why GLM-OCR is aimed at structured documents rather than only at isolated lines of text.
This distinction matters for practical deployments. The base recognition model and the complete document-processing pipeline are not necessarily the same thing. Teams using the open-weight model locally may need to configure layout detection, region handling, preprocessing, batching, and output formatting themselves. The full workflow therefore involves more than downloading model weights and sending an image to a decoder.
Accuracy and throughput
Z.ai reports a score of 94.62 on OmniDocBench V1.5 and states that GLM-OCR ranked first overall at launch. The technical material also reports strong results for text recognition, mathematical formula recognition, table parsing, and key-information extraction. These are provider or project claims and should be interpreted in the context of the benchmark’s documents, metrics, and test conditions rather than as a guarantee for every document type.
Z.ai also reports approximately 1.86 pages per second for PDF documents and 0.67 images per second under its stated test conditions. Actual performance will depend on image resolution, document complexity, hardware, batching, preprocessing, the surrounding layout pipeline, and whether the hosted API or a local deployment is used.
The practical trade-off is straightforward: GLM-OCR gives up general-purpose language and reasoning breadth in exchange for a smaller model focused on document recognition. For a workload dominated by invoices, scanned forms, tables, or PDFs, that specialization can be more useful than paying for a much larger model with capabilities the application does not need.
Supported inputs and output formats
The hosted GLM-OCR service accepts PDF files and common image formats including JPG and PNG. Z.ai documents a maximum single-image size of 10 MB, a maximum PDF size of 50 MB, and support for documents up to 100 pages. These limits apply to the documented hosted workflow and should be checked before submitting large or unusually complex files.
The model supports Chinese, English, French, Spanish, Russian, German, Japanese, Korean, and other languages according to the supplied documentation. Recognition quality can vary with language, font, scan quality, handwriting, page layout, and the amount of visual noise in the source document.
GLM-OCR can return several forms of output:
- Plain text for basic transcription and search indexing.
- Markdown for preserving a readable document structure.
- HTML-formatted table content when table layout needs to be retained.
- Image links for document regions or related visual content.
- Structured JSON for extracting fields such as invoice values, names, dates, identifiers, or other key information.
Its output is text or structured data rather than newly generated images, audio, or video. The model record identifies text output and structured output, but does not establish a separate general-purpose JSON mode with guarantees equivalent to a broad language-model structured-output system.
Pricing and access
The documented hosted API price is $0.03 per million input tokens and $0.03 per million output tokens. This is the listed GLM-OCR API rate, not a subscription price for the wider Z.ai consumer service. The practical cost of a document-processing job will depend on how the service tokenizes the submitted content and generated result.
GLM-OCR is also released as an open-weight model under the MIT License. Z.ai identifies Hugging Face and ModelScope as distribution channels, and the model can be deployed with tools and frameworks including vLLM, SGLang, Ollama, and Transformers. Self-hosting can provide greater control over data handling, concurrency, and infrastructure costs, but it also transfers responsibility for compatible hardware, installation, scaling, monitoring, and the supporting layout-analysis pipeline.
The complete official OCR workflow includes PP-DocLayout-V3. Its licensing and deployment terms should be reviewed separately from the MIT license associated with the GLM-OCR model weights. This is especially important when a production system uses the full pipeline rather than the recognizer alone.
Capabilities and boundaries
GLM-OCR is strongest when the input is a visual document and the desired result is a faithful transcription or structured extraction. It can help process scanned reports, receipts, invoices, certificates, forms, identity documents, screenshots, academic pages, technical material, and handwritten notes. It is also suitable for converting documents into normalized content for search, databases, or retrieval-augmented generation systems.
It is not designed to replace a general-purpose assistant. The supplied research does not describe GLM-OCR as a model for open-ended conversation, advanced reasoning, software development, web search, audio processing, video processing, or image generation. Its reasoning and coding scores are low comparative editorial ratings, not provider-published benchmark results. In practical terms, the model may extract code from a document, but that does not make it a coding assistant; it may recognize a mathematical formula, but that does not establish advanced mathematical reasoning.
The model record lists image input and text output, with no audio or video input and no direct image, audio, or video output. It also lists tool use as unsupported. There is no supplied verified context-window value, maximum output-token value, streaming specification, caching feature, or batch API guarantee. Applications should therefore avoid assuming that GLM-OCR has the same API behavior as a general-purpose language model.
Main strengths and limitations
Where GLM-OCR is strong
- Document specialization: The model targets text, tables, formulas, handwriting, layouts, and key information rather than treating every image as a generic visual question.
- Compact deployment: Its approximately 0.9B parameter scale is suitable for teams seeking a smaller local model or a lower-cost high-volume OCR component.
- Structured results: Markdown, HTML tables, and JSON-style extraction can reduce the work required after recognition.
- Open-weight flexibility: The MIT-licensed model can be downloaded and deployed locally using several supported frameworks.
- Managed and self-hosted options: Teams can use the hosted API for simpler integration or operate their own pipeline when data control and infrastructure customization are priorities.
Limitations to plan for
- Not a general assistant: It is not intended for broad reasoning, coding, conversation, web search, or media generation.
- Pipeline complexity: High-quality document parsing may require layout detection and other components in addition to the base model.
- Document limits: The hosted documentation specifies 10 MB per image, 50 MB per PDF, and up to 100 pages per document.
- Unknown general LLM limits: A context length and maximum output-token limit were not identified in the supplied research.
- Variable real-world accuracy: Benchmark results and reported throughput do not guarantee the same performance on poor scans, unusual layouts, dense handwriting, or specialized forms.
- Operational responsibility when self-hosting: Local deployment requires suitable hardware and management of the layout-analysis and inference stack.
When to choose GLM-OCR
Choose GLM-OCR when the central problem is converting visual documents into reliable text or structured records. It is a particularly plausible choice for invoice and receipt extraction, searchable archives, table recovery, formula transcription, form processing, document ingestion, and OCR services that need to handle many files without deploying a very large vision-language model.
The hosted API is the simpler option when a team wants to avoid operating model infrastructure. The open-weight release is more attractive when documents must remain within a controlled environment, when concurrency needs to be tuned locally, or when the team wants to integrate recognition with its own preprocessing and layout pipeline.
Another type of vision-language model may be more appropriate when the application requires broad visual question answering, multi-step reasoning about an image, conversational interaction, coding assistance, web-connected workflows, or image generation. A larger general-purpose model may also be preferable when OCR is only one small part of a wider assistant. Conversely, a simpler traditional OCR engine may be sufficient for clean, single-column printed text and could be easier to operate if tables, formulas, handwriting, and complex layouts are not important.
Bottom line
GLM-OCR is best understood as a focused document-recognition component rather than a small general-purpose chatbot. Its combination of a compact architecture, layout-aware processing, structured output formats, reported benchmark performance, hosted API, and open-weight deployment makes it relevant for document-heavy applications. The main decision is whether its specialized OCR capability matches the workload: for extracting and organizing information from images and PDFs, it offers a clear speed-and-cost-oriented alternative to larger multimodal models; for general reasoning or interactive assistance, a different model type will be a better fit.

