Hunyuan Vision 1.5

HY-Vision-1.5-Thinking

by Tencent AI · Online and currently available through Tencent Cloud TokenHub

Tencent Hunyuan’s HY-Vision-1.5-Thinking is a multimodal model for reasoning over images and supported video inputs. It handles visual question answering, OCR, chart and document analysis, localization, and educational problem solving, returning text rather than generated media. The TokenHub model ID is hunyuan-t1-vision-20250916, with a 40,000-token context window, 16,000-token maximum input, 24,000-token maximum output, and pricing of CNY 3 per million input tokens and CNY 9 per million output tokens.

Text Reasoning Coding
HY-Vision-1.5-Thinking is a currently available Tencent Hunyuan model built for visual understanding rather than media generation. It accepts text together with visual content, reasons through what appears in images or supported video inputs, and returns text-based answers. Tencent Cloud TokenHub exposes it through the API identifier hunyuan-t1-vision-20250916, with a 40K-token context window and a maximum output of 24K tokens.
Outputs

What HY-Vision-1.5-Thinking can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Streaming Structured output
Model profile

Performance characteristics

8/10 Reasoning
4/10 Coding
6/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Hunyuan Vision 1.5
Model type Multimodal
Context window 40K tokens
Maximum output 24K tokens
Release date 2025-10-06
Status Online and currently available through Tencent Cloud TokenHub
Knowledge cutoff notes

Tencent’s current model catalog and model documentation reviewed for this record do not specify a knowledge cutoff for HY-Vision-1.5-Thinking.

Model notes

The canonical TokenHub model identifier is hunyuan-t1-vision-20250916. Tencent’s official catalog lists a 40K-token context window, 16K maximum input and 24K maximum output. The model is positioned for deep visual reasoning and supports image understanding; Tencent’s HunyuanVision materials also describe image and video understanding capabilities. Release reporting identifies the Hunyuan-Vision-1.5-Thinking announcement on October 6, 2025, while the API identifier contains the September 16, 2025 snapshot date. Tencent does not publicly specify a knowledge cutoff, parameter count, fine-tuning support or batch API support for this exact model in the reviewed documentation. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input CNY 3 per 1M input tokens
Output CNY 9 per 1M output tokens
Model guide

HY-Vision-1.5-Thinking: Tencent Hunyuan’s Model for Visual Reasoning

HY-Vision-1.5-Thinking is Tencent Hunyuan’s multimodal vision-language model for deliberate reasoning over images and other visual inputs. Available through Tencent Cloud TokenHub as hunyuan-t1-vision-20250916, it supports visual question answering, OCR, chart and document analysis, visual localization, educational problem solving, and multilingual image understanding. It offers a 40,000-token context window, up to 16,000 input tokens, and up to 24,000 output tokens. TokenHub pricing is CNY 3 per million input tokens and CNY 9 per million output tokens.

What is HY-Vision-1.5-Thinking?

HY-Vision-1.5-Thinking is Tencent Hunyuan’s multimodal vision-language model for tasks where the answer depends on visual information. In practical terms, it can examine an image, screenshot, chart, diagram, document, or other supported visual input and produce a textual explanation or answer. Its focus is not simply recognizing objects; the model is positioned for deliberate reasoning over visual content.

The model is available through Tencent Cloud TokenHub. Its catalog name is HY-Vision-1.5-Thinking, while its canonical API model identifier is hunyuan-t1-vision-20250916. The model was announced as part of the Hunyuan Vision 1.5 line, and Tencent’s current catalog lists it as online and available through TokenHub.

This positioning makes it different from an image-generation system. HY-Vision-1.5-Thinking analyzes visual inputs and returns text. It is not documented as a model for generating images, video, audio, speech, music, or embeddings.

Where it fits in Tencent’s current catalog

HY-Vision-1.5-Thinking sits in Tencent’s Hunyuan Vision family as a model aimed at visual reasoning and understanding. The “Thinking” designation reflects its intended use for tasks that require more than a short visual description, such as interpreting a chart, locating an item in an image, extracting information from a document, or working through an educational problem shown in a picture.

The supplied Tencent materials describe image understanding as the core use case, while HunyuanVision materials also describe image and video understanding capabilities. The precise behavior available to an application depends on the TokenHub service configuration and the visual inputs included in each request. The model should therefore be selected as a visual-analysis model, not as a general-purpose media model.

Core capabilities and examples

The model’s main strength is connecting visual evidence with a reasoned textual response. Supported use cases described in the supplied research include:

  • Visual question answering: answering questions about the contents, relationships, and meaning of an image.
  • OCR and document understanding: reading text embedded in images and using that text in an explanation or extraction workflow.
  • Chart and diagram analysis: interpreting plotted values, labels, layouts, and relationships in charts or diagrams.
  • Visual localization: identifying where a relevant object, region, or detail appears in an image.
  • Educational problem solving: helping analyze photographed or screenshot-based exercises and explaining a solution in text.
  • Image-grounded dialogue: supporting follow-up questions about the same visual context.
  • Creative image understanding: describing or reasoning about visual content for image-based creative tasks, without generating a new image.
  • Multilingual visual reasoning: handling image question-answering and related analysis across languages, with reported improvements for English and smaller-language scenarios.

For example, an application could submit a screenshot of a dashboard and ask which category has the largest change, upload a scanned form and request field extraction, or provide a diagram and ask for an explanation of how its components relate. These examples illustrate the model’s intended role: the visual input supplies evidence, and the model produces a text answer based on that evidence.

Context window, input limits, and output limits

Tencent’s current model catalog specifies a 40,000-token context window. The catalog separately lists a maximum input length of 16,000 tokens and a maximum output length of 24,000 tokens.

SpecificationVerified value
Context window40,000 tokens
Maximum input16,000 tokens
Maximum output24,000 tokens
API identifierhunyuan-t1-vision-20250916

These figures describe request and response capacity; they are not a statement about the model’s knowledge cutoff. Tencent does not specify a model-specific knowledge cutoff in the reviewed documentation. Applications that depend on recent facts should supply the relevant information in the prompt or through an external retrieval workflow rather than assume that the model has current knowledge.

Supported modalities and output type

HY-Vision-1.5-Thinking supports text input and visual input. The supplied model data identifies image and video input support, although the strongest and most specifically documented use cases are image-based analysis, OCR, charts, documents, and visual question answering.

The output is text. The model does not natively return images, video, audio, speech, music, or embeddings. This distinction matters when designing a workflow: it can inspect a photograph or visual document and explain what it finds, but another model or service is required to create a new image or produce spoken audio.

Reasoning, coding, and tool support

Reasoning over visual evidence is the model’s primary capability. Tencent positions it for deep visual reasoning, multi-turn visual dialogue, visual localization, and problem solving rather than only caption generation. This makes it better suited to questions that require combining several clues in an image or explaining a conclusion step by step.

Coding is not the model’s main specialization. The supplied editorial assessment gives it a coding score of 4 out of 10, compared with a reasoning score of 8 out of 10. These are comparative editorial estimates, not Tencent-published benchmark results. They suggest that HY-Vision-1.5-Thinking is more appropriate for code-related tasks involving screenshots, diagrams, or visual debugging context than for choosing a model solely for software development.

The model data also lists tool use and streaming support. Tool use can allow an application to connect model responses with external actions or services, but the supplied research does not define a complete tool schema or enumerate specific built-in tools. Structured output is listed as supported, while a distinct JSON-mode capability is not verified; developers should confirm the exact response-format behavior in the current TokenHub API documentation.

Pricing and value trade-offs

Current TokenHub pricing is:

  • Input: CNY 3 per million tokens.
  • Output: CNY 9 per million tokens.

Output tokens cost three times as much as input tokens, so long reasoning responses can have a greater effect on cost than short visual prompts. The model’s maximum output of 24,000 tokens is useful for complex explanations, but applications should normally request only the response length they need.

The supplied editorial assessment gives HY-Vision-1.5-Thinking a speed score of 6 out of 10 and a cost score of 7 out of 10. These scores are subjective comparisons rather than provider measurements. They indicate a middle-ground profile: a reasoning-oriented visual model with moderate estimated speed and relatively favorable estimated cost, rather than an option selected primarily for minimum latency or minimum price.

Best use cases

HY-Vision-1.5-Thinking is a strong candidate when the application needs textual reasoning grounded in visual evidence. Suitable workloads include:

  • Extracting text and fields from photographed or scanned documents.
  • Answering questions about screenshots, charts, diagrams, and images.
  • Explaining educational exercises captured by a camera.
  • Building visual customer-support assistants that interpret uploaded images.
  • Locating objects or regions in an image before a downstream workflow acts on them.
  • Classifying, describing, or extracting information from visual search results.
  • Analyzing multilingual image content where ordinary text-only processing is insufficient.
  • Reviewing image or video content and returning a textual summary or interpretation.

It is especially useful when the answer requires both perception and reasoning. A basic OCR service may be enough for plain text extraction, while a general text model may be enough when all relevant information is already transcribed. HY-Vision-1.5-Thinking becomes more attractive when layout, visual relationships, charts, or image context affect the answer.

When to choose this model

Choose HY-Vision-1.5-Thinking when visual understanding is central to the task and the application benefits from a relatively large context and extended textual reasoning. The 40K-token context and 24K-token output ceiling are useful for complex documents, long explanations, or multi-step visual analysis, subject to the separate 16K-token maximum input limit.

It is also a reasonable choice when an OpenAI-compatible TokenHub API surface, streaming, tool integration, and structured responses are useful to the application. Those features can simplify integration, but the exact request and response behavior should be checked against Tencent’s current API documentation.

Another type of model may be more appropriate in several situations. Use an image-generation model when the required result is a new picture rather than an explanation. Use a dedicated speech or audio system for spoken output, and use a text-specialized coding model when software development is the primary workload. A lightweight OCR or classification service may be preferable for very simple, high-volume tasks where deep visual reasoning is unnecessary and minimum latency is more important.

Limitations and unverified specifications

The model’s documented scope has important boundaries. It is a text-output understanding model, not a native media-generation system. Although Tencent materials describe image and video understanding, the supplied research does not provide detailed video duration, frame, file-size, or resolution limits. Those operational limits should be verified before production deployment.

Tencent does not publicly specify a knowledge cutoff, parameter count, fine-tuning support, or batch API support for this exact model in the reviewed materials. Fine-tuning, caching, and batch API availability should therefore be treated as unverified rather than assumed unavailable or supported. The model also does not provide embeddings according to the supplied specification.

Finally, actual cost and performance depend on the amount of visual and textual content sent, tokenization, response length, request parameters, and TokenHub configuration. Editorial scores for reasoning, coding, speed, and cost are useful for broad positioning only; they are not substitutes for testing the model on representative images, documents, charts, and languages used by a particular application.


Answers to Frequently Asked Questions

How much does HY-Vision-1.5-Thinking cost on Tencent Cloud TokenHub?
TokenHub pricing is CNY 3 per million input tokens and CNY 9 per million output tokens. Since output tokens cost three times as much as input tokens, limiting unnecessarily long reasoning responses can help control costs.
What are the context, input, and output limits of HY-Vision-1.5-Thinking?
Tencent’s catalog lists a 40,000-token context window, a maximum input length of 16,000 tokens, and a maximum output length of 24,000 tokens.
What can HY-Vision-1.5-Thinking be used for?
Common use cases include visual question answering, OCR and document understanding, chart and diagram analysis, visual localization, educational problem solving, image-grounded dialogue, multilingual visual reasoning, and textual analysis of image or video content.
What is the API model identifier for HY-Vision-1.5-Thinking?
The catalog name is HY-Vision-1.5-Thinking, and its canonical API model identifier is hunyuan-t1-vision-20250916. The model is available through Tencent Cloud TokenHub.
What is HY-Vision-1.5-Thinking?
HY-Vision-1.5-Thinking is Tencent Hunyuan’s multimodal vision-language model for analyzing images, screenshots, charts, diagrams, documents, and other visual inputs. It produces textual answers and explanations grounded in visual evidence, with an emphasis on deliberate visual reasoning.


Sources 5
Provider

About Tencent AI