Hunyuan turbos-vision

HY-Vision-Video

by Tencent AI · Online and currently listed in Tencent Cloud TokenHub; the older Hunyuan platform entry was retired on 2026-06-22, while the TokenHub model remains available.

Tencent’s HY-Vision-Video is a specialized video-understanding model that accepts video URLs and text instructions, then returns textual descriptions, summaries, answers, and scene or action analysis. It is available through TokenHub with a 32,000-token context window, an 8,000-token maximum output, and pricing of CNY 3 per million input tokens and CNY 9 per million output tokens.

Text Reasoning Coding
HY-Vision-Video is a specialized multimodal model from Tencent’s Hunyuan family. It accepts video and text instructions, then returns textual analysis rather than creating new video. That makes it suited to video question answering, summaries, metadata generation, content review, and other workflows where a system needs to understand what happens inside a video.
Outputs

What HY-Vision-Video can produce

Text
Inputs

What it can understand

Text Video Multimodal input
Model profile

Performance characteristics

3/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Hunyuan turbos-vision
Model type Multimodal
Context window 32K tokens
Maximum output 8K tokens
Release date 2025-07-28
Status Online and currently listed in Tencent Cloud TokenHub; the older Hunyuan platform entry was retired on 2026-06-22, while the TokenHub model remains available.
Knowledge cutoff notes

Tencent does not publish a specific knowledge cutoff for HY-Vision-Video in the cited model catalog, API documentation, pricing documentation, or launch material.

Model notes

The canonical TokenHub model identifier is hunyuan-turbos-vision-video-20250728. Tencent describes the model as a video-understanding model supporting video descriptions and video content question answering. TokenHub uses an OpenAI-compatible Chat Completions endpoint. Video should be supplied through a publicly accessible or pre-signed URL in a video_url content block; Tencent states that HY-Vision-Video does not support Base64 video input on TokenHub. The TokenHub platform limits individual video files and request bodies to 100 MB. Pricing is token-based and may be subject to actual usage, account configuration, and billing changes. Comparative scores are editorial estimates, not vendor benchmarks.

Cost

Model pricing

Input CNY 3 per 1 million input tokens
Output CNY 9 per 1 million output tokens
Model guide

HY-Vision-Video: Tencent’s Model for Video Understanding and Question Answering

HY-Vision-Video is Tencent’s Hunyuan video-understanding model for analyzing uploaded videos, generating descriptions and summaries, answering questions about video content, and supporting scene, action, structure, and content-review workflows through Tencent Cloud TokenHub.

What is HY-Vision-Video?

HY-Vision-Video is Tencent’s dedicated model for understanding video. It is available in Tencent Cloud TokenHub, where its canonical model identifier is hunyuan-turbos-vision-video-20250728. The model belongs to Tencent’s Hunyuan turbos-vision line and is positioned as a video-analysis system, not as a video-generation model.

In practical terms, a user supplies a video and a text instruction such as “Summarize the main events,” “What happens after the person enters the room?” or “List the major scenes and their order.” HY-Vision-Video then produces a text response based on the video and the question. It can therefore serve as an analysis layer for applications that need to search, review, summarize, describe, or organize video libraries.

Tencent’s documentation identifies the model as supporting video descriptions and video-content question answering. The broader uses described for the model include video structure analysis, content review, and action-oriented understanding.

Where it fits in Tencent’s catalog

HY-Vision-Video is a focused member of Tencent’s Hunyuan model family. Its specialization is important: unlike a general text model, it is designed to inspect video; unlike a video-generation model, it does not produce new video clips. Its output is textual analysis.

The model launched on July 28, 2025, under the hunyuan-turbos-vision-video line. It is currently listed as online in Tencent Cloud TokenHub. Tencent also recorded the retirement of an older entry for the same model from a previous Hunyuan platform on June 22, 2026. These are separate availability contexts, so users should use the TokenHub model identifier when configuring the current TokenHub integration.

Supported inputs and outputs

HY-Vision-Video accepts text and video. Video is provided through a video_url content block in an OpenAI-compatible Chat Completions request. TokenHub documentation specifically states that HY-Vision-Video does not support Base64 video input on that platform. A publicly accessible or pre-signed video URL should therefore be used instead.

The model returns text. It does not natively produce video, images, audio, music, embeddings, or other direct non-text media outputs. This distinction matters when choosing between a video-understanding model and a generative media system: HY-Vision-Video can explain what is in an existing video, but it is not intended to create or edit one.

TokenHub imposes a platform-side limit of 100 MB for an individual video and for the request body. The supplied documentation does not establish a fixed maximum video duration, so duration and practical processing behavior may depend on the submitted file and platform implementation. Users should also account for the fact that longer or visually dense videos consume more of the available context and may require focused questions or staged analysis.

Context, output limits, and API access

TokenHub lists a 32,000-token context window for HY-Vision-Video. The listed maximum input allowance is 24,000 tokens, while the maximum output is 8,000 tokens. The context window is the total working space for the request and response; the separate input and output figures describe the documented limits for each side of the interaction.

Access is provided through Tencent Cloud TokenHub using an OpenAI-compatible Chat Completions interface. “OpenAI-compatible” describes the request style and endpoint behavior; it does not mean that every feature of another provider’s API is available. For this model, the supplied research does not verify function calling, tool use, JSON mode, structured outputs, streaming, caching, batch processing, or fine-tuning. Those capabilities should be treated as unconfirmed rather than assumed.

What HY-Vision-Video does well

The model’s main strength is its narrow focus on video understanding. It is designed to connect visual events with natural-language instructions, which is useful when a simple caption is not enough. For example, a workflow can ask it to identify the sequence of events, describe changes between scenes, answer a question about an action, or extract information needed for a review process.

  • Video descriptions: Generate a readable account of what a video shows.
  • Video question answering: Answer targeted questions about people, actions, events, or scenes in the supplied footage.
  • Summarization: Produce concise overviews for longer recordings or collections of clips.
  • Structure analysis: Organize a video around scenes, events, or changes in activity.
  • Content review: Support human review and classification workflows by identifying relevant video content.
  • Metadata generation: Create descriptions, labels, or searchable text for video libraries.
  • Action analysis: Help identify and explain actions or events that occur in the footage.

These uses are consistent with Tencent’s stated positioning. They should not be read as a guarantee of perfect recognition, exhaustive event detection, or suitability for high-stakes decisions without human verification.

Pricing and cost profile

TokenHub lists pricing of CNY 3 per million input tokens and CNY 9 per million output tokens. Input billing covers the material sent to the model, while output billing covers the generated text. Actual charges can depend on usage, account configuration, and later changes to Tencent Cloud pricing.

On the supplied editorial scoring scale, HY-Vision-Video receives a speed score of 8 and a cost score of 8. These are editorial estimates, not Tencent benchmarks. They indicate that the model may be attractive for teams prioritizing relatively efficient video-analysis calls and token-based pricing, but they do not establish a guaranteed response time or total cost per video. Video size, request frequency, prompt length, and the amount of requested explanation will all affect practical usage.

Reasoning, coding, and tool capabilities

HY-Vision-Video is primarily an interpretation model. Its useful reasoning behavior is tied to understanding and explaining video content rather than solving broad mathematical, research, or planning problems. The supplied editorial reasoning score is 3, which should be treated as a subjective assessment of its general reasoning profile, not as a provider-published benchmark.

Coding is not a target use case. The supplied editorial coding score is 1, and Tencent’s documented positioning focuses on video descriptions and questions rather than software development. The model may be able to follow simple text instructions around an analysis task, but it should not be selected as a general coding assistant.

Tool use is listed as unsupported in the supplied model record. No verified function-calling or external-tool capability is provided for this model. If an application needs automated database actions, web search, or structured agent operations, those functions would need to be implemented outside HY-Vision-Video and connected around its text response, subject to the capabilities of the surrounding platform.

Best use cases

HY-Vision-Video is a sensible choice when the central input is an existing video and the desired result is a written explanation. Suitable examples include:

  • Creating searchable summaries for recorded meetings, training sessions, or media archives.
  • Answering user questions about the contents of an uploaded clip.
  • Generating initial descriptions and metadata for a video-management system.
  • Reviewing footage for scenes, actions, or events that deserve human attention.
  • Building a question-and-answer interface over a collection of videos.
  • Producing structured notes about the order and content of scenes.

For a large archive, teams may benefit from asking focused questions and storing the resulting descriptions or metadata rather than repeatedly sending the same videos for broad analysis. The 24,000-token input allowance and 100 MB request constraint should be considered when designing ingestion and review pipelines.

When should you choose HY-Vision-Video?

Choose HY-Vision-Video when video understanding is the primary requirement, the output can be text, and TokenHub access is appropriate for your deployment. Its specialization makes it more relevant than a text-only model for questions that depend on visual events. Its listed CNY 3 per million input-token and CNY 9 per million output-token rates may also suit workloads that need repeated analysis and do not require media generation.

A different option may be more appropriate in several situations. Use a video-generation system when the goal is to create or transform footage rather than interpret it. Use an image, audio, or speech model when the key input or output is a different media type. Use a general reasoning or coding model for complex programming, broad analytical tasks, or tool-driven workflows. If an application requires verified structured output, function calling, streaming, fine-tuning, or batch processing, those features should be confirmed in the current TokenHub documentation before selecting HY-Vision-Video.

Limitations to plan for

The most important limitation is the model’s scope: it understands supplied video and responds with text, but it is not a general-purpose media creation system. Video must be delivered through a supported URL rather than Base64 on TokenHub, and each video and request body is subject to the 100 MB platform limit.

The supplied research does not specify a knowledge cutoff, fixed video-duration limit, benchmark accuracy, guaranteed latency, or detailed behavior for difficult visual conditions. It also does not verify tool use, structured output, streaming, caching, batch API access, or fine-tuning. Those unknowns should be tested against the current service documentation and the application’s own data before production deployment.

Overall, HY-Vision-Video is best understood as a focused Tencent model for turning video into useful text. Its value comes from video-specific understanding, question answering, and analysis—not from general coding, agent tools, or media generation.


Answers to Frequently Asked Questions

Can HY-Vision-Video generate videos, use tools, or write code?
No. HY-Vision-Video is designed to interpret video and produce text, not generate or edit media. Tool use and function calling are not verified for this model, and coding is not a target use case. Applications requiring these capabilities should use other models or implement them externally.
What are the main limits and pricing for HY-Vision-Video?
TokenHub lists a 32,000-token context window, with a maximum of 24,000 input tokens and 8,000 output tokens. Each video and request body is limited to 100 MB. Pricing is listed as CNY 3 per million input tokens and CNY 9 per million output tokens.
What is HY-Vision-Video used for?
HY-Vision-Video is Tencent’s video-understanding model for analyzing existing video and returning text. It can describe footage, summarize events, answer questions about people or actions, analyze scene structure, and generate searchable metadata.
How do I provide video to HY-Vision-Video through TokenHub?
Use an OpenAI-compatible Chat Completions request with the video supplied in a video_url content block. TokenHub does not support Base64 video input for HY-Vision-Video, so the video should be available through a publicly accessible or pre-signed URL.


Sources 6
Provider

About Tencent AI