HY-Vision

HY-Vision-2.0-Instruct

by Tencent AI · Current and available through Tencent Cloud TokenHub

HY-Vision-2.0-Instruct is Tencent’s current TokenHub model for text-based image understanding. It supports text and multiple-image inputs for OCR, visual question answering, chart and diagram interpretation, STEM reasoning, image description, and comparison. The model has a 44,000-token context window, 24,000-token maximum input, 16,000-token maximum output, and pricing of ¥7.5 per million input tokens and ¥17.5 per million output tokens.

Text Reasoning Coding
Tencent HY-Vision-2.0-Instruct is a fast-thinking multimodal model for image-to-text work. Available through Tencent Cloud TokenHub, it combines textual instructions with image inputs to answer questions, extract text, interpret charts, compare images, and reason about visual content. The model offers a 44K-token context window and usage-based pricing of ¥7.5 per million input tokens and ¥17.5 per million output tokens.
Outputs

What HY-Vision-2.0-Instruct can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

7/10 Reasoning
4/10 Coding
8/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family HY-Vision
Model type Multimodal
Context window 44K tokens
Maximum output 16K tokens
Status Current and available through Tencent Cloud TokenHub
Knowledge cutoff notes

Tencent's current public model documentation does not state a knowledge-cutoff date for HY-Vision-2.0-Instruct.

Model notes

The canonical API model identifier is hy-vision-2.0-instruct. Tencent's current TokenHub catalog lists a 44K-token context window, a 24K maximum input, and a 16K maximum output. The model is described as a fast-thinking multimodal understanding model for general image-to-text scenarios, with improvements in perception, content recognition, knowledge, OCR, STEM, reasoning, and chart understanding. Tencent documentation demonstrates single-image and multi-image requests using an OpenAI-compatible chat-completions API. Pricing is usage-based and listed in Chinese yuan per million tokens. No authoritative exact-model evidence was found for a separate legacy JSON mode, fine-tuning, batch API, audio input, video input, native non-text output, or web-search grounding.

Cost

Model pricing

Input ¥7.5 per 1 million input tokens
Output ¥17.5 per 1 million output tokens
Model guide

HY-Vision-2.0-Instruct: Tencent’s Fast Image-Understanding Model

HY-Vision-2.0-Instruct is Tencent’s current multimodal image-understanding model on Tencent Cloud TokenHub. It accepts text and one or more images and returns text for visual question answering, image description, OCR, chart interpretation, STEM reasoning, and related analysis. Its documented 44,000-token context window, 24,000-token maximum input, and 16,000-token maximum output make it suitable for detailed visual tasks, while its text-only output means it is not an image, audio, or video generation model.

What is HY-Vision-2.0-Instruct?

HY-Vision-2.0-Instruct is Tencent’s current multimodal understanding model for visual analysis. In practical terms, it reads images alongside a written prompt and produces a written answer. A request might ask it to describe a photograph, read text from a document, explain a chart, compare several images, or solve a visual STEM problem.

The model is listed in Tencent Cloud TokenHub’s multimodal understanding catalog. Its canonical API identifier is hy-vision-2.0-instruct. Tencent describes it as a fast-thinking model and reports improvements over the previous generation in basic perception, content recognition, knowledge, optical character recognition (OCR), STEM reasoning, inference, and chart understanding.

Those improvement statements are provider claims rather than independent benchmark results. The supplied documentation establishes the model’s intended capabilities and published limits, but does not provide a separate benchmark table for this exact model.

Where it fits in Tencent’s catalog

HY-Vision-2.0-Instruct is positioned as an image-understanding model, not as a general-purpose image generator or a model for producing audio and video. Its role is to turn visual information into useful textual analysis. This makes it relevant to applications that need to inspect existing visual material rather than create new media.

The model can be used for both simple perception and more involved reasoning. For example, a basic request could ask what objects appear in an image, while a more demanding request could ask it to interpret a graph, identify the relationship between multiple diagrams, or answer a question based on a screenshot and accompanying instructions.

Supported inputs and outputs

HY-Vision-2.0-Instruct accepts text and images. Tencent’s documentation demonstrates image content supplied through URLs or Base64 data URLs in an OpenAI-compatible chat-completions interface. The model supports both single-image and multiple-image requests.

CapabilityDocumented status
Text inputSupported
Image inputSupported
Multiple-image requestsSupported
Audio inputNot established in the supplied documentation
Video inputNot established in the supplied documentation
Text outputSupported
Image, audio, or video outputNot documented; the model is intended for textual responses

The distinction between multimodal input and multimodal output is important. HY-Vision-2.0-Instruct can process visual material, but its documented response is text. It should therefore not be selected when an application needs the model to generate or edit images, synthesize speech, or produce video.

What the model does well

Visual question answering and description

The model can answer questions about image content and produce general descriptions. This covers common tasks such as identifying visible objects, summarizing a scene, explaining a screenshot, or answering a question whose evidence is contained in an image.

OCR and document reading

OCR allows a model to recognize written content in an image. HY-Vision-2.0-Instruct is documented as supporting OCR and content recognition, making it relevant to scanned pages, photographed documents, screenshots, signs, and other image-based text. The model’s output is still generated text, so applications that require exact extraction should validate important fields rather than assuming every character will be transcribed perfectly.

Charts, diagrams, and STEM reasoning

Tencent specifically identifies chart understanding and STEM reasoning among the model’s capabilities. This makes it useful for explaining trends in a chart, locating values, interpreting visual relationships, or answering questions about diagrams and other technical images. These tasks combine perception with reasoning: the model must first identify relevant visual elements and then relate them to the user’s question.

Multi-image comparison

Multiple-image requests allow an application to provide more than one visual input in a single task. Potential uses include comparing document pages, examining before-and-after images, matching diagrams with accompanying screenshots, or asking the model to identify differences between pictures. Tencent documents multi-image usage, although the supplied research does not state a fixed maximum image count. The practical limit is related to the available context window and the size of the supplied content.

Context window and output limits

The published context window is 44,000 tokens. Tencent lists a maximum input of 24,000 tokens and a maximum output of 16,000 tokens. A token is a unit used to measure processed text; in a multimodal request, the overall context also needs to accommodate the visual inputs and the surrounding prompt.

LimitPublished valuePractical meaning
Total context window44,000 tokensThe request and response must fit within the model’s available context.
Maximum input24,000 tokensSets the documented upper bound for supplied prompt and input content.
Maximum output16,000 tokensAllows substantially detailed textual answers when the task requires them.

These limits are useful for document and multi-image workflows, but a large theoretical context does not mean every request should use the maximum. Smaller prompts and focused questions can reduce cost and make responses easier to evaluate. When sending several images, applications should also account for how the images are represented and consumed by the service.

Pricing and API access

Tencent Cloud TokenHub currently lists usage-based pricing of ¥7.5 per million input tokens and ¥17.5 per million output tokens. Input and output are priced separately, so the final cost depends on both the amount of visual or textual material sent to the model and the length of its response.

Usage typePrice
Input¥7.5 per million tokens
Output¥17.5 per million tokens

These are TokenHub usage prices expressed in Chinese yuan, not a subscription fee or a flat monthly plan. Account access, regional availability, quotas, and billing requirements should be confirmed with Tencent before production deployment because those operational details can affect whether the listed endpoint is usable for a particular application.

The model is accessed through an OpenAI-compatible chat-completions interface. This compatibility describes the request style documented by Tencent; it does not by itself establish support for every feature associated with other platforms. The supplied research does not verify native function calling, tool use, streaming, fine-tuning, batch processing, caching, or a separate legacy JSON mode for this exact model.

Reasoning, coding, speed, and cost trade-offs

HY-Vision-2.0-Instruct is primarily a visual reasoning model. Its useful reasoning work is tied to images: interpreting charts, answering questions from diagrams, extracting meaning from documents, and comparing visual inputs. It is not documented as a specialist software-coding model.

For editorial comparison purposes, the supplied evaluation rates its reasoning at 7 out of 10, coding at 4 out of 10, speed at 8 out of 10, and cost at 6 out of 10. These are subjective editorial scores, not Tencent-published benchmarks or guarantees. The relatively stronger speed assessment reflects the model’s stated fast-thinking positioning, while the lower coding assessment reflects the lack of evidence that it is optimized for standalone programming tasks.

The pricing creates a practical trade-off. At ¥7.5 per million input tokens, repeated visual inspection can be economical when requests are concise, while output costs ¥17.5 per million tokens. Applications should avoid asking for unnecessarily long explanations when a short extraction or classification is sufficient. Conversely, the 16K output allowance can support detailed reports when the additional response length is genuinely useful.

Best use cases

  • Document and screenshot analysis: Ask questions about pages, interfaces, forms, or captured workflows.
  • OCR-assisted extraction: Convert text in images into a preliminary structured or prose representation for downstream checking.
  • Chart and diagram interpretation: Explain trends, relationships, labels, and visual evidence in technical or business graphics.
  • Visual question answering: Respond to questions about photographs, illustrations, screenshots, and other images.
  • STEM visual reasoning: Analyze diagrams and solve questions that require both visual perception and reasoning.
  • Multi-image comparison: Compare related pages, images, or diagrams within one request.

For production systems, important OCR fields and conclusions should be checked with application-level validation or human review. Visual models can produce plausible explanations even when an image is ambiguous, low quality, cropped, or difficult to read.

When to choose HY-Vision-2.0-Instruct

Choose HY-Vision-2.0-Instruct when the central requirement is text-based understanding of one or more images and you want a model positioned for fast responses. It is a reasonable fit for visual support tools, document triage, chart explanations, screenshot assistants, and image-based question answering.

Another type of model may be more appropriate when the task requires image generation or editing, audio or video processing, embeddings, or a verified tool-calling workflow. A text-focused coding model is likely a better choice for software generation that does not depend on visual input. Likewise, an application requiring guaranteed structured output should not assume that the OpenAI-compatible interface provides a distinct JSON mode, because that capability is not verified for this exact model in the supplied documentation.

HY-Vision-2.0-Instruct is also not the best choice when the required input modality is audio or video, since those capabilities are not established. In those cases, use a service whose documentation explicitly supports the needed modality rather than treating general multimodal branding as evidence of support.

Limitations and verification points

The model’s documented scope is image understanding with textual responses. There is no supplied evidence of image generation, native audio or video input, non-text output, embeddings, fine-tuning, batch API access, caching, web-search grounding, or standalone JSON-mode support. These omissions should be treated as unverified rather than as confirmed impossibilities if Tencent later expands the service.

Before relying on the model in a production workflow, verify the current TokenHub model listing, regional availability, quotas, authentication requirements, pricing, and request limits. Also test representative images from the intended workload. A model that performs well on clear charts may behave differently on small text, unusual layouts, poor lighting, dense documents, or images requiring highly precise numerical extraction.

Overall, HY-Vision-2.0-Instruct is best understood as a current Tencent TokenHub model for fast, text-producing visual analysis. Its strongest distinction is the combination of broad image-understanding tasks, multi-image support, and a 44K-token context window—not the ability to generate media or act as a general-purpose coding and tool-use model.


Answers to Frequently Asked Questions

How much does HY-Vision-2.0-Instruct cost on Tencent Cloud TokenHub?
Tencent Cloud TokenHub lists usage-based pricing of ¥7.5 per million input tokens and ¥17.5 per million output tokens. Input and output are billed separately, so the total cost depends on the amount of supplied content and the length of the generated response.
When should I choose HY-Vision-2.0-Instruct?
Choose HY-Vision-2.0-Instruct when an application needs fast, text-based analysis of one or more images, such as document triage, OCR-assisted extraction, chart interpretation, screenshot assistance, visual question answering, or multi-image comparison. It is not the best choice for image generation, audio or video processing, standalone coding, or unverified tool-calling and JSON-mode workflows.
What are the context window and token limits for HY-Vision-2.0-Instruct?
HY-Vision-2.0-Instruct has a published 44,000-token context window, with a maximum input of 24,000 tokens and a maximum output of 16,000 tokens. Applications must account for the prompt, response, and visual inputs when staying within these limits.
What is HY-Vision-2.0-Instruct used for?
HY-Vision-2.0-Instruct is Tencent’s multimodal image-understanding model for analyzing images alongside text prompts. It can describe images, answer visual questions, read documents with OCR, interpret charts and diagrams, solve visual STEM problems, and compare multiple images.
What inputs and outputs does HY-Vision-2.0-Instruct support?
The model supports text and image inputs, including single-image and multiple-image requests. Images can be supplied through URLs or Base64 data URLs in Tencent’s OpenAI-compatible chat-completions interface. Its documented output is text; audio, video, image generation, and other non-text outputs are not established.


Sources 5
Provider

About Tencent AI