What is HunyuanOCR-1.5?
HunyuanOCR-1.5 is Tencent’s current open-weight model for optical character recognition (OCR) and document understanding. OCR normally means detecting and transcribing text in an image, but this model is intended to do more than return a plain text string. It can interpret the layout and content of a document and produce formats suited to downstream software, including Markdown, HTML tables, LaTeX formulas, translations, and structured extraction results.
The model has approximately 1 billion parameters, making it substantially smaller and more deployment-oriented than many general-purpose vision-language models. Its design prioritizes text-rich images, scanned documents, forms, tables, charts, formulas, and mixed-language pages. The current official model repository is tencent/HunyuanOCR; the active checkpoint is HunyuanOCR-1.5, while the earlier HunyuanOCR-1.0 checkpoint remains archived in the repository’s v1.0 directory.
Tencent released HunyuanOCR-1.5 on July 7, 2026. The 1.5 release keeps the lightweight architecture but adds an upgraded training recipe, Agentic Data Flow data construction, DFlash speculative decoding for faster long-form generation, and PC-side deployment through llama.cpp.
What can HunyuanOCR-1.5 do?
HunyuanOCR-1.5 combines visual input with task instructions. Instead of treating every page as a simple image-to-text problem, users can request a particular type of result. The official inference toolkit exposes task types including doc_parse, structured_parse, spotting_json, layout_parse, chart_parse, formula, and table.
- Document parsing: Convert pages into readable Markdown or other text-based representations while preserving useful document structure.
- Structured parsing: Extract fields from forms, reports, invoices, and other documents into a structured result.
- Text spotting: Identify text within an image and return localization-oriented information, including JSON-oriented spotting workflows.
- Layout parsing: Distinguish document regions and help preserve relationships between headings, paragraphs, tables, figures, and other elements.
- Table recognition: Convert visual tables into structured or HTML-like representations instead of flattening them into an unreadable sequence of words.
- Formula recognition: Transcribe mathematical formulas into LaTeX or another text representation.
- Chart parsing: Process charts and their associated labels or data-oriented content.
- Text-image translation: Translate text contained in visual documents, including multilingual or mixed-language pages.
These capabilities make the model most useful when the document’s arrangement matters. For example, extracting a table, preserving a formula, or identifying fields in a form requires more than recognizing individual characters.
Inputs, outputs, and supported modalities
The model accepts text instructions and images. In practical use, an image is supplied with a prompt specifying the desired operation, such as parsing the page, extracting a table, or returning text-spotting coordinates. The model’s outputs are text or structured text. It does not natively generate images, audio, or video.
This distinction is important when evaluating HunyuanOCR-1.5 as a multimodal model. It has multimodal input because it can process visual documents and text prompts, but its output is text-centered. Its structured results may be formatted as JSON, Markdown, HTML, or LaTeX depending on the task, but the supplied research does not identify a separate general-purpose JSON-mode guarantee. Structured parsing and JSON-oriented task workflows should therefore be treated as task-specific output formats rather than proof of a universal constrained-decoding mode.
Deployment, speed, and local use
HunyuanOCR-1.5 is primarily distributed as downloadable weights for self-hosted inference under the Tencent Hunyuan Community License Agreement. The documented deployment options include Transformers, vLLM, SGLang, and llama.cpp after conversion to GGUF. This gives developers several deployment paths: Python-based model execution, production-oriented inference servers, or local CPU/GPU-oriented use through llama.cpp.
The 1.5 release adds DFlash speculative decoding. Speculative decoding uses a faster auxiliary process to help accelerate generation from the main model, which is particularly relevant for OCR jobs that produce long structured documents. Tencent positions this addition as a way to improve long-form inference speed. The supplied research does not provide a universal tokens-per-second benchmark, so actual performance will vary with hardware, image resolution, quantization, serving framework, batching, and output length.
The model is also marked as supporting streaming and batch inference in the supplied technical record. Batch inference is provided by the open-source toolkit rather than a separately documented Tencent-hosted batch API. This distinction matters: local batch processing can be useful for document archives, but it does not imply that Tencent operates a managed batch service for this checkpoint.
Context and output limits
Tencent’s vLLM setup documents a maximum model length of 131,072 tokens. This is the relevant published context-length figure for the serving configuration, but it should not be interpreted as a guarantee that every deployment can process that much content efficiently. Image resolution, visual complexity, GPU memory, batching, and the selected inference framework all affect practical capacity.
The official client supports generation requests of up to 32,768 tokens for long OCR outputs. That limit concerns generated output rather than the complete model context. A large document can therefore still be constrained by the relationship between input images, prompts, context usage, and the requested response length. For ordinary single-page OCR, the maximum is unlikely to be the limiting factor; it becomes more relevant for multi-page parsing and detailed structured extraction.
Pricing and access
No official Tencent hosted per-token price was identified for HunyuanOCR-1.5. The model’s primary distribution method is downloadable open-weight files for self-hosted inference, so the direct model price is not a conventional monthly or per-token API subscription.
Self-hosting is not cost-free. Users must provide suitable compute, storage, and operational infrastructure, and the total cost depends on whether the model runs on a workstation, a consumer device, or a production server. A smaller model can reduce inference requirements compared with larger vision-language systems, but the research does not establish a universal hardware requirement or cost figure. Developers should also review the Tencent Hunyuan Community License Agreement before incorporating the checkpoint into a commercial workflow.
Main strengths and trade-offs
The clearest strength of HunyuanOCR-1.5 is specialization. It targets document text, layout, tables, formulas, charts, and extraction rather than trying to be a general conversational assistant. That focus can make it a better fit for repeatable OCR pipelines where the output format and document structure matter.
Its approximately 1-billion-parameter size is another practical advantage for local deployment. The combination of open weights, several serving options, GGUF support through llama.cpp, streaming, and batch tooling gives teams more control over data handling and infrastructure than a hosted-only service. DFlash speculative decoding is specifically relevant to workloads that generate long document representations.
There are also clear trade-offs. A specialized OCR model is not the right default for broad reasoning, open-ended dialogue, software development, image generation, audio processing, or video generation. The supplied editorial assessment rates its reasoning and coding suitability below its speed and cost suitability; those are editorial evaluations, not Tencent-published benchmark scores. It should not be presented as a general-purpose replacement for a larger multimodal reasoning model.
Tool or function calling is not documented for this checkpoint in the supplied research. Likewise, no general web-search capability is identified. Applications that need browsing, external actions, or agent-style orchestration must provide those capabilities around the model rather than assuming they are built in.
Best use cases
- Converting scanned documents into Markdown or other structured text.
- Extracting tables from reports, forms, and other visual documents.
- Recognizing formulas and preserving them as LaTeX.
- Parsing charts, layouts, and document regions.
- Extracting named fields from documents using structured parsing tasks.
- Text spotting and localization-oriented OCR workflows.
- Processing multilingual or mixed-language documents.
- Running OCR locally when documents should remain within a controlled environment.
- Batch-processing document collections with an open-source inference toolkit.
When to choose HunyuanOCR-1.5
Choose HunyuanOCR-1.5 when the central problem is visual text understanding and you want downloadable weights, local deployment, and task-specific document outputs. It is particularly suitable when tables, formulas, layout, or structured fields are more important than conversational interaction. The model’s size and supported deployment paths also make it worth considering when inference speed, infrastructure control, or operating cost matters more than maximum general-purpose reasoning ability.
A hosted OCR API may be more appropriate when a team does not want to manage GPUs, model serving, upgrades, or license review. A larger general-purpose vision-language model may be preferable when OCR is only one part of a workflow that also requires complex reasoning, coding, broad dialogue, tool use, or interpretation of non-document imagery. A conventional OCR engine may remain preferable for a narrowly defined, high-volume text transcription task if advanced layout, chart, formula, or extraction capabilities are unnecessary.
Limitations to consider
HunyuanOCR-1.5 requires a compatible vision-language serving stack and adequate compute for efficient inference. Practical results will depend on image quality, resolution, page complexity, language mix, serving configuration, and the requested output format. The published materials do not provide a general knowledge cutoff, and that concept is less relevant to a model specialized in visual OCR than to a conversational language model.
The model also has no identified official hosted price for the exact checkpoint, no documented generic tool-use interface, and no separate general JSON-mode guarantee. Its structured task types should not be confused with broad agent capabilities. For applications requiring guaranteed schemas, external validation may still be appropriate.
Overall, HunyuanOCR-1.5 is best understood as a focused, locally deployable document intelligence model. Its value comes from combining OCR with layout-aware parsing and structured document output, not from serving as an all-purpose multimodal assistant.

