VITA

YT-VITA

by Tencent AI · Currently accessible; scheduled for shutdown on October 15, 2026 at 00:00 Beijing time on Tencent Cloud TokenHub and ADP

Tencent YT-VITA is a hosted multimodal understanding model that analyzes images, videos, audio, and text and returns textual results. It is designed for video summarization, structural parsing, tagging, classification, timestamp extraction, and object localization. TokenHub documents a 128k context window, 100k maximum input, 15k maximum output, and pricing of ¥1.2 per million input tokens plus ¥3.5 per million output tokens. Video must be supplied by URL, and TokenHub and ADP availability is scheduled to end on October 15, 2026.

Text Reasoning Coding
Tencent YT-VITA is designed for applications that need to understand multimedia rather than generate it. Through Tencent Cloud TokenHub, the model can analyze images, videos, audio, and text, then return textual descriptions, summaries, classifications, timestamps, and other requested analysis. Its 128,000-token context window and 15,000-token maximum output make it suitable for substantial media-understanding tasks, although video input must be supplied by URL through TokenHub. The most important practical consideration is availability: Tencent has announced that the TokenHub and ADP YT-VITA service will be removed at 00:00 Beijing time on October 15, 2026.
Outputs

What YT-VITA can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

5/10 Reasoning
3/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family VITA
Model type Multimodal
Context window 128K tokens
Maximum output 15K tokens
Release date 2026-06-11
Status Currently accessible; scheduled for shutdown on October 15, 2026 at 00:00 Beijing time on Tencent Cloud TokenHub and ADP
Shutdown date 2026-10-15
Model notes

YT-VITA is the TokenHub model identity for Tencent's Youtu-VITA service and uses the API identifier youtu-vita. Tencent documentation describes support for image, video, audio, and text understanding with textual responses. TokenHub lists a 128k context window, 100k maximum input, and 15k maximum output. Video input should be supplied by URL in TokenHub because Base64 video input is not supported for YT-VITA. Tencent announced that YT-VITA will be removed from TokenHub and ADP at 00:00 Beijing time on October 15, 2026. Tencent recommends HY-Vision models as migration targets; a separate visual-understanding interface may continue offering a VITA version.

Cost

Model pricing

Input ¥1.2 per million tokens
Output ¥3.5 per million tokens
Model guide

Tencent YT-VITA: Multimodal Video Understanding with a Scheduled TokenHub Shutdown

Tencent YT-VITA is a hosted multimodal understanding model for analyzing images, video, audio, and text. Available through Tencent Cloud TokenHub under the API identifier youtu-vita, it focuses on video structure parsing, summarization, classification, tagging, object localization, and related enterprise media workflows. It accepts multimodal input but produces text, and its TokenHub and ADP availability is scheduled to end on October 15, 2026.

What is Tencent YT-VITA?

YT-VITA is Tencent Youtu Lab's multimodal understanding model, exposed through Tencent Cloud TokenHub with the API model identifier youtu-vita. It is built to interpret multimedia content and respond with text. In practical terms, an application can provide a video, image, audio input, or text prompt and ask the model to describe, classify, summarize, locate, or structurally analyze what it contains.

The model belongs to the understanding side of the AI market, not the media-generation side. It does not produce images, videos, music, or speech as its native output. Its value is in turning visual and audio content into usable textual information, such as a video summary, a list of scenes, content labels, object locations, or a timestamped explanation.

YT-VITA is therefore aimed at workflows where separate computer-vision, speech-recognition, and language-processing stages might otherwise have to be assembled. The supplied research describes it as a native multimodal model, although the exact capabilities available can depend on the Tencent interface and media type being used.

Where YT-VITA fits in Tencent's catalog

YT-VITA is currently offered through Tencent Cloud TokenHub and related Tencent Cloud ADP integrations. TokenHub presents it as a model that can be accessed through an OpenAI-compatible chat-completions interface. This positioning makes YT-VITA a hosted service for developers and enterprise teams that want to integrate multimedia analysis into an application rather than deploy a model locally.

Tencent's documentation and launch material position the model around visual and multimedia understanding. Tencent recommends HY-Vision models as migration targets for users affected by the announced TokenHub and ADP shutdown. That recommendation is relevant for service continuity, but it does not make HY-Vision the subject of this page or establish that it has identical capabilities.

The announced shutdown is a significant part of YT-VITA's current product position. According to the supplied Tencent announcement, YT-VITA will be removed from TokenHub and ADP at 00:00 Beijing time on October 15, 2026. A separate visual-understanding interface may continue to provide a VITA version with matching pricing and model behavior, but users should verify that availability and interface independently before planning a new production deployment.

Supported inputs and output

YT-VITA accepts text, images, video, and audio according to the supplied model specifications. Text prompts can guide the model toward a particular output structure or analytical goal. For example, a prompt might request a short summary, a classification label, a list of detected objects, or a breakdown of scenes and timestamps.

  • Images: image descriptions, classification, tagging, visual question answering, and object localization.
  • Video: summaries, shot or scene analysis, structural decomposition, timestamp extraction, and content review.
  • Audio: direct analysis and summarization of speech or other supported audio content.
  • Text: instructions that define the task, categories, questions, or desired response format.

The model's output is text. This distinction matters because multimodal input does not imply multimodal generation. YT-VITA can analyze a video or image, but it is not the appropriate choice when the application needs a newly generated image, video, audio recording, or song.

For TokenHub video requests, the supplied documentation states that video should be passed by URL because Base64 video input is not supported for YT-VITA. This is an implementation constraint that should be checked early: a media pipeline that only exposes local Base64 data may need to add a secure URL-based storage or delivery step.

Context window and technical limits

TokenHub lists a 128,000-token context window, a 100,000-token maximum input, and a 15,000-token maximum output. The context window represents the total amount of information the service can consider within a request context, while the maximum input and output values describe the documented limits for the respective portions of the exchange.

SpecificationDocumented value
ProviderTencent
Model identifieryoutu-vita
Model typeMultimodal understanding
Context window128,000 tokens
Maximum input100,000 tokens
Maximum output15,000 tokens
Input modalitiesText, image, video, and audio
Output modalityText
StreamingSupported

These figures are provider-documented service limits, not guarantees that every video or audio file will consume tokens in the same way. The effective request size can depend on the media type, the endpoint, and how the content is represented. Developers should also confirm current quotas, request-size rules, and URL-access requirements in the endpoint they intend to use.

YT-VITA's main strengths

YT-VITA's strongest feature is the combination of multiple input types in one understanding workflow. A media-review application can ask about visual content, audio, and written instructions without necessarily maintaining separate model calls for each modality. This can simplify workflows such as cataloging a video library, summarizing meetings or podcasts, and reviewing livestream content.

Its video-oriented capabilities are especially relevant. The supplied research describes structural video parsing, shot or scene breakdown, timestamp extraction, summarization, tagging, and object localization. These functions can support a searchable media archive, automated content operations, e-commerce video review, or moderation queues where operators need a concise explanation before deciding what to do next.

The 128,000-token context window is also useful for long analytical prompts or substantial multimodal tasks, subject to the media-specific restrictions of the service. Streaming support can help applications display partial textual results while a longer analysis is being returned, although streaming does not change the model's underlying output type or maximum output limit.

Reasoning, coding, and tool support

YT-VITA is primarily a media-understanding model rather than a general-purpose reasoning or coding model. The supplied editorial evaluation assigns it a reasoning score of 5 out of 10 and a coding score of 3 out of 10. These are editorial scores for comparison, not Tencent-published benchmark results or official capability ratings.

The model can follow text-guided instructions about how to analyze media, but the research does not describe it as a dedicated reasoning model. Similarly, it may be able to return structured explanations or code-like text when prompted, but it is not documented here as a coding specialist. Teams choosing a model for software development, complex code generation, or general analytical reasoning should evaluate a model designed for those tasks instead.

The supplied specifications mark tool or function support as unavailable. YT-VITA should therefore be treated as a model that analyzes the supplied content and returns text, not as an agent that independently calls external tools, browses the web, or performs application actions. Structured output or JSON mode is also not verified in the supplied research, so applications that require strict machine-readable responses should confirm current endpoint behavior rather than assume it.

Pricing and availability

Current TokenHub pricing lists YT-VITA at ¥1.2 per million input tokens and ¥3.5 per million output tokens. These are separate input and output rates, so the final cost depends on how much content the request sends and how much text the model returns. A long video-analysis instruction or a detailed response can affect the bill differently because the two token categories have different prices.

Tencent's launch material also advertises a one-million-token trial allowance for new accounts. Trial eligibility, account conditions, quotas, and validity should be confirmed with Tencent before being included in a cost forecast.

Availability requires special care. The supplied research states that YT-VITA remains accessible through Tencent Cloud TokenHub and related ADP integrations as of September 25, 2026, but that the TokenHub and ADP service is scheduled for removal at 00:00 Beijing time on October 15, 2026. This makes the listed price useful for current evaluation and short-term operation, but insufficient by itself for a long-lived production decision.

Speed and cost trade-offs

The supplied editorial evaluation gives YT-VITA a speed score of 7 out of 10 and a cost score of 8 out of 10. These scores are subjective comparisons, not provider guarantees. The relatively favorable cost assessment is connected to the documented per-million-token rates, while actual spend and response time will depend on media size, request complexity, traffic, and service conditions.

YT-VITA may be more economical than building a multi-stage workflow for every multimedia request, because one model can handle image, video, audio, and text understanding. That does not mean it is automatically the cheapest option for every job. A simple image-labeling task may be better served by a specialized vision service, while a pure transcription task may be better suited to a dedicated speech-recognition system. YT-VITA is most compelling when the application needs several kinds of interpretation together, especially for video.

Best use cases for YT-VITA

  • Video summarization: create concise descriptions of long videos for users who need an overview before watching.
  • Video structure analysis: identify scenes, shots, transitions, or important timestamps for indexing and navigation.
  • Media tagging and classification: assign labels to image and video collections for search, organization, or review.
  • Object localization: identify objects and explain where they appear in visual content.
  • Livestream and e-commerce analysis: review product demonstrations, promotional footage, or live content at scale.
  • Meeting and podcast summaries: analyze supported audio or video recordings and produce textual summaries.
  • Enterprise multimedia search: turn large media libraries into text-based metadata that can support discovery and content operations.

Limitations and when another option may be better

YT-VITA is not suitable for native media generation. Choose an image, video, speech, or music generation model when the task is to create new media rather than interpret existing content. It is also not the best fit for general coding, advanced software engineering, or tool-using agent workflows because those capabilities are not documented as core features here.

Input handling is another limitation. TokenHub video requests require a URL rather than Base64 video, so organizations must account for media hosting, access control, and URL availability. The exact limits for different media types may also depend on the interface, even though the general TokenHub specification lists 128,000 context tokens, 100,000 maximum input tokens, and 15,000 maximum output tokens.

Finally, the scheduled shutdown changes the model-selection calculation. YT-VITA can make sense for an existing TokenHub workload that needs immediate multimodal understanding, a favorable documented token price, and a short-term evaluation path. It is a riskier choice for a new system expected to run unchanged beyond October 15, 2026. Those projects should assess Tencent's recommended HY-Vision migration path or another currently supported service, while verifying that the replacement meets the required video, audio, output, pricing, and availability requirements.

Bottom line

Tencent YT-VITA is a focused multimodal understanding model for turning images, video, audio, and text into textual analysis. Its practical strengths are video summarization, structural parsing, tagging, classification, timestamp-oriented review, and object localization, supported by a large documented context window and TokenHub access. Its weaknesses are equally clear: text-only output, no verified tool support, limited suitability for coding or general reasoning, URL-based video input requirements, and a scheduled TokenHub and ADP shutdown.

For current multimedia analysis, YT-VITA offers a coherent alternative to stitching together several specialized services. For a new long-term deployment, however, service continuity should be treated as a first-order requirement rather than an afterthought.


Answers to Frequently Asked Questions

When will YT-VITA be removed from Tencent Cloud TokenHub and ADP?
According to the supplied Tencent announcement, YT-VITA will be removed from TokenHub and ADP at 00:00 Beijing time on October 15, 2026. Tencent recommends HY-Vision models as migration targets, but users should independently verify replacement capabilities, pricing, interface compatibility, and availability before moving a production workload.
How much does YT-VITA cost on Tencent Cloud TokenHub?
The listed TokenHub pricing is ¥1.2 per million input tokens and ¥3.5 per million output tokens. Tencent launch material also mentions a one-million-token trial allowance for new accounts, but eligibility, quotas, and validity should be confirmed with Tencent.
What are YT-VITA's documented token limits?
Tencent Cloud TokenHub lists a 128,000-token context window, a 100,000-token maximum input, and a 15,000-token maximum output for YT-VITA. Actual media-related limits and token usage can vary by endpoint, media type, and content representation, so developers should verify the current service documentation.
What input and output modalities does YT-VITA support?
YT-VITA accepts text, images, video, and audio as inputs and produces text as its output. It is designed to interpret existing media rather than generate new images, videos, music, or speech. For TokenHub video requests, video must be provided by URL because Base64 video input is not supported.
What is Tencent YT-VITA used for?
Tencent YT-VITA is a multimodal understanding model for analyzing text, images, video, and audio and returning textual results. Common uses include video summarization, scene and shot analysis, timestamp extraction, media tagging, classification, object localization, meeting summaries, podcast analysis, and enterprise multimedia search.


Sources 7
Provider

About Tencent AI