What is HY-Vision-Video?
HY-Vision-Video is Tencent’s dedicated model for understanding video. It is available in Tencent Cloud TokenHub, where its canonical model identifier is hunyuan-turbos-vision-video-20250728. The model belongs to Tencent’s Hunyuan turbos-vision line and is positioned as a video-analysis system, not as a video-generation model.
In practical terms, a user supplies a video and a text instruction such as “Summarize the main events,” “What happens after the person enters the room?” or “List the major scenes and their order.” HY-Vision-Video then produces a text response based on the video and the question. It can therefore serve as an analysis layer for applications that need to search, review, summarize, describe, or organize video libraries.
Tencent’s documentation identifies the model as supporting video descriptions and video-content question answering. The broader uses described for the model include video structure analysis, content review, and action-oriented understanding.
Where it fits in Tencent’s catalog
HY-Vision-Video is a focused member of Tencent’s Hunyuan model family. Its specialization is important: unlike a general text model, it is designed to inspect video; unlike a video-generation model, it does not produce new video clips. Its output is textual analysis.
The model launched on July 28, 2025, under the hunyuan-turbos-vision-video line. It is currently listed as online in Tencent Cloud TokenHub. Tencent also recorded the retirement of an older entry for the same model from a previous Hunyuan platform on June 22, 2026. These are separate availability contexts, so users should use the TokenHub model identifier when configuring the current TokenHub integration.
Supported inputs and outputs
HY-Vision-Video accepts text and video. Video is provided through a video_url content block in an OpenAI-compatible Chat Completions request. TokenHub documentation specifically states that HY-Vision-Video does not support Base64 video input on that platform. A publicly accessible or pre-signed video URL should therefore be used instead.
The model returns text. It does not natively produce video, images, audio, music, embeddings, or other direct non-text media outputs. This distinction matters when choosing between a video-understanding model and a generative media system: HY-Vision-Video can explain what is in an existing video, but it is not intended to create or edit one.
TokenHub imposes a platform-side limit of 100 MB for an individual video and for the request body. The supplied documentation does not establish a fixed maximum video duration, so duration and practical processing behavior may depend on the submitted file and platform implementation. Users should also account for the fact that longer or visually dense videos consume more of the available context and may require focused questions or staged analysis.
Context, output limits, and API access
TokenHub lists a 32,000-token context window for HY-Vision-Video. The listed maximum input allowance is 24,000 tokens, while the maximum output is 8,000 tokens. The context window is the total working space for the request and response; the separate input and output figures describe the documented limits for each side of the interaction.
Access is provided through Tencent Cloud TokenHub using an OpenAI-compatible Chat Completions interface. “OpenAI-compatible” describes the request style and endpoint behavior; it does not mean that every feature of another provider’s API is available. For this model, the supplied research does not verify function calling, tool use, JSON mode, structured outputs, streaming, caching, batch processing, or fine-tuning. Those capabilities should be treated as unconfirmed rather than assumed.
What HY-Vision-Video does well
The model’s main strength is its narrow focus on video understanding. It is designed to connect visual events with natural-language instructions, which is useful when a simple caption is not enough. For example, a workflow can ask it to identify the sequence of events, describe changes between scenes, answer a question about an action, or extract information needed for a review process.
- Video descriptions: Generate a readable account of what a video shows.
- Video question answering: Answer targeted questions about people, actions, events, or scenes in the supplied footage.
- Summarization: Produce concise overviews for longer recordings or collections of clips.
- Structure analysis: Organize a video around scenes, events, or changes in activity.
- Content review: Support human review and classification workflows by identifying relevant video content.
- Metadata generation: Create descriptions, labels, or searchable text for video libraries.
- Action analysis: Help identify and explain actions or events that occur in the footage.
These uses are consistent with Tencent’s stated positioning. They should not be read as a guarantee of perfect recognition, exhaustive event detection, or suitability for high-stakes decisions without human verification.
Pricing and cost profile
TokenHub lists pricing of CNY 3 per million input tokens and CNY 9 per million output tokens. Input billing covers the material sent to the model, while output billing covers the generated text. Actual charges can depend on usage, account configuration, and later changes to Tencent Cloud pricing.
On the supplied editorial scoring scale, HY-Vision-Video receives a speed score of 8 and a cost score of 8. These are editorial estimates, not Tencent benchmarks. They indicate that the model may be attractive for teams prioritizing relatively efficient video-analysis calls and token-based pricing, but they do not establish a guaranteed response time or total cost per video. Video size, request frequency, prompt length, and the amount of requested explanation will all affect practical usage.
Reasoning, coding, and tool capabilities
HY-Vision-Video is primarily an interpretation model. Its useful reasoning behavior is tied to understanding and explaining video content rather than solving broad mathematical, research, or planning problems. The supplied editorial reasoning score is 3, which should be treated as a subjective assessment of its general reasoning profile, not as a provider-published benchmark.
Coding is not a target use case. The supplied editorial coding score is 1, and Tencent’s documented positioning focuses on video descriptions and questions rather than software development. The model may be able to follow simple text instructions around an analysis task, but it should not be selected as a general coding assistant.
Tool use is listed as unsupported in the supplied model record. No verified function-calling or external-tool capability is provided for this model. If an application needs automated database actions, web search, or structured agent operations, those functions would need to be implemented outside HY-Vision-Video and connected around its text response, subject to the capabilities of the surrounding platform.
Best use cases
HY-Vision-Video is a sensible choice when the central input is an existing video and the desired result is a written explanation. Suitable examples include:
- Creating searchable summaries for recorded meetings, training sessions, or media archives.
- Answering user questions about the contents of an uploaded clip.
- Generating initial descriptions and metadata for a video-management system.
- Reviewing footage for scenes, actions, or events that deserve human attention.
- Building a question-and-answer interface over a collection of videos.
- Producing structured notes about the order and content of scenes.
For a large archive, teams may benefit from asking focused questions and storing the resulting descriptions or metadata rather than repeatedly sending the same videos for broad analysis. The 24,000-token input allowance and 100 MB request constraint should be considered when designing ingestion and review pipelines.
When should you choose HY-Vision-Video?
Choose HY-Vision-Video when video understanding is the primary requirement, the output can be text, and TokenHub access is appropriate for your deployment. Its specialization makes it more relevant than a text-only model for questions that depend on visual events. Its listed CNY 3 per million input-token and CNY 9 per million output-token rates may also suit workloads that need repeated analysis and do not require media generation.
A different option may be more appropriate in several situations. Use a video-generation system when the goal is to create or transform footage rather than interpret it. Use an image, audio, or speech model when the key input or output is a different media type. Use a general reasoning or coding model for complex programming, broad analytical tasks, or tool-driven workflows. If an application requires verified structured output, function calling, streaming, fine-tuning, or batch processing, those features should be confirmed in the current TokenHub documentation before selecting HY-Vision-Video.
Limitations to plan for
The most important limitation is the model’s scope: it understands supplied video and responds with text, but it is not a general-purpose media creation system. Video must be delivered through a supported URL rather than Base64 on TokenHub, and each video and request body is subject to the 100 MB platform limit.
The supplied research does not specify a knowledge cutoff, fixed video-duration limit, benchmark accuracy, guaranteed latency, or detailed behavior for difficult visual conditions. It also does not verify tool use, structured output, streaming, caching, batch API access, or fine-tuning. Those unknowns should be tested against the current service documentation and the application’s own data before production deployment.
Overall, HY-Vision-Video is best understood as a focused Tencent model for turning video into useful text. Its value comes from video-specific understanding, question answering, and analysis—not from general coding, agent tools, or media generation.

