HunyuanOCR

HunyuanOCR-1.5

by Tencent AI · Current open-weight release

Tencent’s HunyuanOCR-1.5 is an approximately 1-billion-parameter open-weight vision-language model specialized in OCR and document understanding. It handles document and layout parsing, text spotting, structured extraction, tables, formulas, charts, translation, and multilingual visual text. The model supports Transformers, vLLM, SGLang, and llama.cpp deployment, with a documented 131,072-token vLLM model length and client generation requests up to 32,768 tokens. No official hosted per-token price was identified.

Text Reasoning Coding
HunyuanOCR-1.5 is a lightweight OCR-focused vision-language model from Tencent Hunyuan. It reads text and visual documents, then produces results such as Markdown, HTML tables, LaTeX formulas, translations, or structured extraction data. Compared with a general multimodal chatbot, it is designed for a narrower but highly practical purpose: converting complex visual documents into usable text and structured information. The open-weight checkpoint can be deployed with Transformers, vLLM, SGLang, or llama.cpp, although Tencent does not publish a hosted per-token price for this exact model.
Outputs

What HunyuanOCR-1.5 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming Fine-tuning Structured output Batch API
Model profile

Performance characteristics

3/10 Reasoning
2/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family HunyuanOCR
Model type Multimodal
Context window 131K tokens
Maximum output 33K tokens
Release date 2026-07-07
Status Current open-weight release
Knowledge cutoff notes

Tencent's official model materials do not publish a knowledge cutoff for HunyuanOCR-1.5. The model is specialized for visual OCR and document understanding, so a general conversational knowledge-cutoff interpretation may not be applicable.

Model notes

The current official repository and Hugging Face model identifier are tencent/HunyuanOCR, but the active checkpoint is HunyuanOCR-1.5. HunyuanOCR-1.0 is archived under the v1.0 subdirectory. The model is approximately 1B parameters and is distributed under the Tencent Hunyuan Community License Agreement. Official materials document a 131,072-token vLLM model length and client generation requests up to 32,768 tokens; practical limits depend on serving configuration and hardware. The model supports document parsing, structured parsing, text spotting, layout, charts, formulas, tables, and translation workflows. DFlash speculative decoding accelerates long autoregressive outputs. Batch inference is provided by the open-source toolkit, not as a separately documented Tencent hosted batch API. No official hosted per-token price or knowledge cutoff was identified for this exact checkpoint.

Model guide

HunyuanOCR-1.5: Tencent’s Lightweight Model for Document OCR and Structured Extraction

HunyuanOCR-1.5 is Tencent’s open-weight, approximately 1-billion-parameter vision-language model for OCR, document parsing, text spotting, information extraction, tables, formulas, charts, translation, and other text-centered image tasks. Its 1.5 release adds DFlash speculative decoding, improved long-output handling, multilingual coverage, and llama.cpp deployment while remaining focused on self-hosted document understanding rather than general-purpose conversation.

What is HunyuanOCR-1.5?

HunyuanOCR-1.5 is Tencent’s current open-weight model for optical character recognition (OCR) and document understanding. OCR normally means detecting and transcribing text in an image, but this model is intended to do more than return a plain text string. It can interpret the layout and content of a document and produce formats suited to downstream software, including Markdown, HTML tables, LaTeX formulas, translations, and structured extraction results.

The model has approximately 1 billion parameters, making it substantially smaller and more deployment-oriented than many general-purpose vision-language models. Its design prioritizes text-rich images, scanned documents, forms, tables, charts, formulas, and mixed-language pages. The current official model repository is tencent/HunyuanOCR; the active checkpoint is HunyuanOCR-1.5, while the earlier HunyuanOCR-1.0 checkpoint remains archived in the repository’s v1.0 directory.

Tencent released HunyuanOCR-1.5 on July 7, 2026. The 1.5 release keeps the lightweight architecture but adds an upgraded training recipe, Agentic Data Flow data construction, DFlash speculative decoding for faster long-form generation, and PC-side deployment through llama.cpp.

What can HunyuanOCR-1.5 do?

HunyuanOCR-1.5 combines visual input with task instructions. Instead of treating every page as a simple image-to-text problem, users can request a particular type of result. The official inference toolkit exposes task types including doc_parse, structured_parse, spotting_json, layout_parse, chart_parse, formula, and table.

  • Document parsing: Convert pages into readable Markdown or other text-based representations while preserving useful document structure.
  • Structured parsing: Extract fields from forms, reports, invoices, and other documents into a structured result.
  • Text spotting: Identify text within an image and return localization-oriented information, including JSON-oriented spotting workflows.
  • Layout parsing: Distinguish document regions and help preserve relationships between headings, paragraphs, tables, figures, and other elements.
  • Table recognition: Convert visual tables into structured or HTML-like representations instead of flattening them into an unreadable sequence of words.
  • Formula recognition: Transcribe mathematical formulas into LaTeX or another text representation.
  • Chart parsing: Process charts and their associated labels or data-oriented content.
  • Text-image translation: Translate text contained in visual documents, including multilingual or mixed-language pages.

These capabilities make the model most useful when the document’s arrangement matters. For example, extracting a table, preserving a formula, or identifying fields in a form requires more than recognizing individual characters.

Inputs, outputs, and supported modalities

The model accepts text instructions and images. In practical use, an image is supplied with a prompt specifying the desired operation, such as parsing the page, extracting a table, or returning text-spotting coordinates. The model’s outputs are text or structured text. It does not natively generate images, audio, or video.

This distinction is important when evaluating HunyuanOCR-1.5 as a multimodal model. It has multimodal input because it can process visual documents and text prompts, but its output is text-centered. Its structured results may be formatted as JSON, Markdown, HTML, or LaTeX depending on the task, but the supplied research does not identify a separate general-purpose JSON-mode guarantee. Structured parsing and JSON-oriented task workflows should therefore be treated as task-specific output formats rather than proof of a universal constrained-decoding mode.

Deployment, speed, and local use

HunyuanOCR-1.5 is primarily distributed as downloadable weights for self-hosted inference under the Tencent Hunyuan Community License Agreement. The documented deployment options include Transformers, vLLM, SGLang, and llama.cpp after conversion to GGUF. This gives developers several deployment paths: Python-based model execution, production-oriented inference servers, or local CPU/GPU-oriented use through llama.cpp.

The 1.5 release adds DFlash speculative decoding. Speculative decoding uses a faster auxiliary process to help accelerate generation from the main model, which is particularly relevant for OCR jobs that produce long structured documents. Tencent positions this addition as a way to improve long-form inference speed. The supplied research does not provide a universal tokens-per-second benchmark, so actual performance will vary with hardware, image resolution, quantization, serving framework, batching, and output length.

The model is also marked as supporting streaming and batch inference in the supplied technical record. Batch inference is provided by the open-source toolkit rather than a separately documented Tencent-hosted batch API. This distinction matters: local batch processing can be useful for document archives, but it does not imply that Tencent operates a managed batch service for this checkpoint.

Context and output limits

Tencent’s vLLM setup documents a maximum model length of 131,072 tokens. This is the relevant published context-length figure for the serving configuration, but it should not be interpreted as a guarantee that every deployment can process that much content efficiently. Image resolution, visual complexity, GPU memory, batching, and the selected inference framework all affect practical capacity.

The official client supports generation requests of up to 32,768 tokens for long OCR outputs. That limit concerns generated output rather than the complete model context. A large document can therefore still be constrained by the relationship between input images, prompts, context usage, and the requested response length. For ordinary single-page OCR, the maximum is unlikely to be the limiting factor; it becomes more relevant for multi-page parsing and detailed structured extraction.

Pricing and access

No official Tencent hosted per-token price was identified for HunyuanOCR-1.5. The model’s primary distribution method is downloadable open-weight files for self-hosted inference, so the direct model price is not a conventional monthly or per-token API subscription.

Self-hosting is not cost-free. Users must provide suitable compute, storage, and operational infrastructure, and the total cost depends on whether the model runs on a workstation, a consumer device, or a production server. A smaller model can reduce inference requirements compared with larger vision-language systems, but the research does not establish a universal hardware requirement or cost figure. Developers should also review the Tencent Hunyuan Community License Agreement before incorporating the checkpoint into a commercial workflow.

Main strengths and trade-offs

The clearest strength of HunyuanOCR-1.5 is specialization. It targets document text, layout, tables, formulas, charts, and extraction rather than trying to be a general conversational assistant. That focus can make it a better fit for repeatable OCR pipelines where the output format and document structure matter.

Its approximately 1-billion-parameter size is another practical advantage for local deployment. The combination of open weights, several serving options, GGUF support through llama.cpp, streaming, and batch tooling gives teams more control over data handling and infrastructure than a hosted-only service. DFlash speculative decoding is specifically relevant to workloads that generate long document representations.

There are also clear trade-offs. A specialized OCR model is not the right default for broad reasoning, open-ended dialogue, software development, image generation, audio processing, or video generation. The supplied editorial assessment rates its reasoning and coding suitability below its speed and cost suitability; those are editorial evaluations, not Tencent-published benchmark scores. It should not be presented as a general-purpose replacement for a larger multimodal reasoning model.

Tool or function calling is not documented for this checkpoint in the supplied research. Likewise, no general web-search capability is identified. Applications that need browsing, external actions, or agent-style orchestration must provide those capabilities around the model rather than assuming they are built in.

Best use cases

  • Converting scanned documents into Markdown or other structured text.
  • Extracting tables from reports, forms, and other visual documents.
  • Recognizing formulas and preserving them as LaTeX.
  • Parsing charts, layouts, and document regions.
  • Extracting named fields from documents using structured parsing tasks.
  • Text spotting and localization-oriented OCR workflows.
  • Processing multilingual or mixed-language documents.
  • Running OCR locally when documents should remain within a controlled environment.
  • Batch-processing document collections with an open-source inference toolkit.

When to choose HunyuanOCR-1.5

Choose HunyuanOCR-1.5 when the central problem is visual text understanding and you want downloadable weights, local deployment, and task-specific document outputs. It is particularly suitable when tables, formulas, layout, or structured fields are more important than conversational interaction. The model’s size and supported deployment paths also make it worth considering when inference speed, infrastructure control, or operating cost matters more than maximum general-purpose reasoning ability.

A hosted OCR API may be more appropriate when a team does not want to manage GPUs, model serving, upgrades, or license review. A larger general-purpose vision-language model may be preferable when OCR is only one part of a workflow that also requires complex reasoning, coding, broad dialogue, tool use, or interpretation of non-document imagery. A conventional OCR engine may remain preferable for a narrowly defined, high-volume text transcription task if advanced layout, chart, formula, or extraction capabilities are unnecessary.

Limitations to consider

HunyuanOCR-1.5 requires a compatible vision-language serving stack and adequate compute for efficient inference. Practical results will depend on image quality, resolution, page complexity, language mix, serving configuration, and the requested output format. The published materials do not provide a general knowledge cutoff, and that concept is less relevant to a model specialized in visual OCR than to a conversational language model.

The model also has no identified official hosted price for the exact checkpoint, no documented generic tool-use interface, and no separate general JSON-mode guarantee. Its structured task types should not be confused with broad agent capabilities. For applications requiring guaranteed schemas, external validation may still be appropriate.

Overall, HunyuanOCR-1.5 is best understood as a focused, locally deployable document intelligence model. Its value comes from combining OCR with layout-aware parsing and structured document output, not from serving as an all-purpose multimodal assistant.


Answers to Frequently Asked Questions

When should developers choose HunyuanOCR-1.5 instead of a general-purpose vision-language model?
HunyuanOCR-1.5 is a strong fit when the main requirement is local, structured document processing involving tables, formulas, layouts, charts, forms, or text extraction. A larger general-purpose vision-language model may be preferable for complex reasoning, coding, broad dialogue, tool use, or non-document image understanding.
Does HunyuanOCR-1.5 have an official hosted API price?
No official Tencent hosted per-token price was identified for HunyuanOCR-1.5. Its primary access model is self-hosting, so users must account for compute, storage, infrastructure, and operational costs, as well as review the Tencent Hunyuan Community License Agreement for commercial use.
How can HunyuanOCR-1.5 be deployed locally?
HunyuanOCR-1.5 is distributed primarily as downloadable weights under the Tencent Hunyuan Community License Agreement. It can be deployed with Transformers, vLLM, SGLang, or llama.cpp after conversion to GGUF, supporting local CPU or GPU use, streaming, and batch processing through the open-source toolkit.
What is HunyuanOCR-1.5?
HunyuanOCR-1.5 is Tencent’s approximately 1-billion-parameter open-weight model for optical character recognition and document understanding. It can process text-rich images, scanned documents, forms, tables, charts, formulas, and multilingual pages, producing outputs such as Markdown, HTML, LaTeX, translations, and structured extraction results.
What tasks does HunyuanOCR-1.5 support?
Its documented task types include document parsing, structured parsing, text spotting, layout parsing, chart parsing, formula recognition, and table recognition. It can preserve document structure, extract fields, identify text locations, convert tables into structured formats, and transcribe mathematical formulas into LaTeX.


Sources 4
Provider

About Tencent AI