WAND-Dubbing-Clone

WAND-Dubbing-Clone-V1

by Tencent AI · Available

Tencent’s WAND-Dubbing-Clone-V1 is an asynchronous video-localization model. It accepts a video URL and language settings, translates speech and subtitles, generates voice-preserving dubbed audio, and returns a translated video after task-based processing. The article explains its supported modalities, workflow, resolution-dependent pricing, strengths, limitations, and fit compared with Tencent’s newer WAND-Dubbing-Clone-V2.

Video generation Speech Reasoning Coding
WAND-Dubbing-Clone-V1 is a specialized Tencent Cloud model for translating and dubbing video. A workflow sends the model a video URL together with source-language and target-language settings; Tencent then processes the job asynchronously and returns a task identifier for status checks. Once processing is complete, the result is a translated video with cloned-voice dubbing and translated subtitles. This makes the model relevant to online courses, films, and short-form video localization, but not to ordinary chat, software development, real-time voice conversations, or standalone text-to-speech tasks.
Outputs

What WAND-Dubbing-Clone-V1 can produce

Video generation Speech
Inputs

What it can understand

Text Audio Video Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
5/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family WAND-Dubbing-Clone
Model type Other
Status Available
Model notes

The canonical model identifier is wand-dubbing-clone-v1. Tencent Cloud documents it as an asynchronous video-translation and dubbing model. The workflow accepts a video URL and source and target language parameters, then returns a task ID for status polling and a translated video result. Tencent's documentation identifies WAND-Dubbing-Clone-V2 as the newer version with more natural wording and higher-fidelity preservation of original emotion, tone, and prosody. Listed pricing is resolution-dependent: 12,280 tokens per second for 720p and below, 14,780 tokens per second for 1080p, and 19,780 tokens per second for 2K and above, at CNY 10 per million tokens.

Cost

Model pricing

Input CNY 10 per 1 million tokens; usage is resolution-dependent and billed according to video duration and token consumption.
Model guide

WAND-Dubbing-Clone-V1: Tencent’s Asynchronous Video Translation and Voice-Dubbing Model

WAND-Dubbing-Clone-V1 is a Tencent Cloud model for asynchronous video localization. It translates spoken content and subtitles into a target language, creates voice-preserving dubbed audio, and returns a translated video. Its main advantage is an end-to-end dubbing workflow rather than general text generation, while its main limitations are asynchronous processing, resolution-dependent token billing, and the absence of documented general-purpose reasoning, coding, or tool-use features.

What is WAND-Dubbing-Clone-V1?

WAND-Dubbing-Clone-V1 is an audiovisual translation model provided through Tencent Cloud. Its purpose is to localize an existing video for viewers who speak another language. Instead of producing only a text translation or a separate audio file, the documented workflow is designed to produce a translated video with dubbed speech and translated subtitles.

The model is built for asynchronous processing. In practical terms, an application submits a video URL and language parameters, receives a task ID, and later polls for the task status and result. This is different from a live conversational system, where audio is sent and returned continuously with low latency. WAND-Dubbing-Clone-V1 is better understood as a media-processing service for completed or uploaded video assets.

Tencent documents wand-dubbing-clone-v1 as the canonical model identifier. The model is listed as available, but the supplied documentation does not provide a release date, context-window size, maximum output-token limit, or deprecation date.

How the dubbing workflow works

The documented process has three practical stages:

  1. Submit the source video: provide a video URL and specify the relevant source and target languages.
  2. Wait for asynchronous processing: Tencent returns a task ID rather than requiring the caller to keep a live audio stream open.
  3. Retrieve the result: poll the task status until the translated video is ready.

The model handles the central localization tasks: translating speech, translating subtitles, and generating dubbed speech that is intended to preserve the source speaker’s voice characteristics. The supplied research describes this as voice-cloned dubbing. The exact supported language list, file-duration limits, URL requirements, and processing-time targets are not specified in the available material, so those details should be verified in Tencent Cloud’s current service documentation before implementation.

Main capabilities and modalities

WAND-Dubbing-Clone-V1 accepts video as its primary input and uses audio from that video as part of the translation and dubbing workflow. Text is also relevant to the request because source and target language parameters and subtitle content are involved, but this is not a general text-generation model with a conventional text-in/text-out interface.

CapabilityWhat is verified
Video inputSupported; the documented workflow accepts a video URL.
Audio inputSupported as part of the source video’s speech and dubbing workflow.
Video outputSupported; the result is a translated video.
Audio or speech outputSupported through dubbed speech in the translated video.
Subtitle translationSupported according to the supplied Tencent documentation summary.
StreamingNot documented as supported; the workflow is asynchronous.
Tool or function callingNot documented as supported.
Structured JSON outputNot documented as a model capability.

The output should therefore be evaluated as a finished media-localization result, not as a response to be consumed like a chatbot message. The model’s direct non-text output is central to its purpose: it produces translated video and dubbed speech rather than merely describing how a video could be translated.

Where it fits in Tencent’s catalog

WAND-Dubbing-Clone-V1 occupies a specialized position in Tencent’s catalog. It is not presented as a general-purpose language model, image generator, coding assistant, or real-time speech agent. Its focus is the complete translation and dubbing of video content.

Tencent also documents WAND-Dubbing-Clone-V2 as a newer version. The supplied documentation says that V2 provides more natural wording and higher-fidelity preservation of the original emotion, tone, and prosody. Those are provider-documented positioning claims about the newer version, not independent benchmark results. They suggest that teams starting a new localization project should compare V1 with V2 where the newer model is available, especially when natural delivery and expressive voice performance matter.

V1 may still be relevant when a project specifically requires this model identifier, has an existing integration, or finds its pricing and output quality suitable. However, the available research does not establish whether V2 costs more, whether it supports exactly the same workflow, or whether V1 has a scheduled shutdown date.

Where WAND-Dubbing-Clone-V1 is strong

  • End-to-end video localization: the service combines translation, subtitle handling, voice dubbing, and translated-video production in one workflow.
  • Voice-preserving dubbing: the model is designed to generate cloned-voice dubbing rather than replacing every speaker with an unrelated generic voice.
  • Useful for asynchronous production: a task-based workflow is appropriate for courses, films, marketing libraries, and short-form content that can be processed before publication.
  • Resolution-aware billing: Tencent’s pricing system reflects the amount of video processing through token consumption that varies by resolution.
  • Clear specialization: teams do not need to assemble a separate general translator, speech synthesizer, and video post-processing pipeline for the basic documented use case.

These strengths are most valuable when the desired deliverable is a localized video. If the required output is only a translated transcript, a text translation model will usually be a more direct choice. If the requirement is live interpretation or interactive voice conversation, an asynchronous video job is a poor fit regardless of translation quality.

Limitations and documented unknowns

The most important operational limitation is asynchronous execution. The supplied research does not describe streaming input or streaming output, so the model should not be selected for applications that require immediate, turn-by-turn responses. A production system also needs to account for task polling, retry handling, failed jobs, and the time required to download or distribute the resulting video.

Several conventional model specifications are not available in the supplied documentation. Tencent does not provide a context length or maximum output-token value for this item, and those concepts are less directly applicable to a video-localization job than they are to a chat model. No release date, deprecation date, or shutdown date is supplied either. The available material also does not establish a detailed language matrix, maximum video duration, supported codecs, subtitle file formats, or guaranteed turnaround time.

WAND-Dubbing-Clone-V1 should not be treated as a reasoning or coding model. The internal evaluation data assigns it a reasoning score of 2 and a coding score of 1, but these are editorial or catalog evaluations, not Tencent-published benchmark results. They indicate that reasoning and coding are outside the model’s intended role. Similarly, the catalog assigns a speed score of 5 and a cost score of 7; these are comparative editorial scores, not provider guarantees.

Pricing and token consumption

Tencent’s documented rate is CNY 10 per 1 million tokens. For this model, token consumption is tied to video processing and resolution rather than being presented as a simple per-request fee. The supplied pricing documentation lists the following consumption rates:

Video resolutionDocumented consumption
720p and below12,280 tokens per second
1080p14,780 tokens per second
2K and above19,780 tokens per second

At the stated rate, higher-resolution video consumes more tokens for every second processed. For example, a 1080p source is billed using the 14,780-token-per-second tier, while a 720p-or-lower source uses the 12,280-token-per-second tier. The exact invoice can depend on Tencent Cloud’s billing rules, rounding, and the measured duration, so these figures should be used for estimation rather than treated as a complete billing calculator.

The pricing structure creates a direct quality-versus-cost decision. Higher resolution may be important for films or premium educational content, but lower-resolution material can reduce processing consumption when the final distribution format does not require 1080p or 2K output.

When to choose WAND-Dubbing-Clone-V1

Choose WAND-Dubbing-Clone-V1 when the main deliverable is a translated and dubbed video and asynchronous processing is acceptable. It is a practical candidate for:

  • localizing online courses for international learners;
  • translating films or recorded programs for regional distribution;
  • creating multilingual versions of short-form videos;
  • preserving a recognizable speaker voice across translated versions; and
  • processing a library of existing videos through a task-based production queue.

It is less appropriate for general-purpose writing, code generation, document question answering, standalone translation of text, live interpretation, real-time conversational speech, image generation, or producing speech from text without a video-translation task. Those use cases call for a model or service designed around the relevant input and latency requirements.

Teams should also compare the model with Tencent’s newer WAND-Dubbing-Clone-V2 when expressive delivery, natural wording, or stronger preservation of emotion and prosody is a priority. V2 is the documented newer sibling, but the available research does not include a direct price or benchmark comparison. The final choice should therefore consider actual sample outputs, language coverage, availability, and current Tencent Cloud billing.

Overall assessment

WAND-Dubbing-Clone-V1 is best viewed as a specialized asynchronous video-localization service rather than a conventional AI model for conversation. Its distinguishing value is the combination of speech translation, subtitle translation, voice-cloned dubbing, and translated-video output. That combination can simplify multilingual media production, particularly for courses, films, and short videos.

Its trade-offs are equally clear: processing is not presented as real time, pricing scales with video duration and resolution, and the supplied documentation does not publish many of the limits commonly associated with text models. For a new project, the strongest reason to select V1 is a need for its documented video-dubbing workflow and available model identifier. For projects beginning from scratch, Tencent’s documented WAND-Dubbing-Clone-V2 should also be evaluated as the newer option, while simpler translation or speech services may be more efficient when a complete translated video is not required.


Answers to Frequently Asked Questions

Is WAND-Dubbing-Clone-V1 suitable for real-time translation or live conversations?
No. WAND-Dubbing-Clone-V1 is designed for asynchronous video processing, not streaming or turn-by-turn interaction. It is better suited to courses, films, recorded programs, marketing libraries, and other completed video assets.
How much does WAND-Dubbing-Clone-V1 cost?
Tencent documents a rate of CNY 10 per 1 million tokens. Consumption depends on video resolution: 12,280 tokens per second for 720p and below, 14,780 tokens per second for 1080p, and 19,780 tokens per second for 2K and above.
What is WAND-Dubbing-Clone-V1 used for?
WAND-Dubbing-Clone-V1 is an asynchronous Tencent Cloud service for localizing videos. It translates spoken content and subtitles, generates voice-cloned dubbed speech, and produces a translated video.
How does the WAND-Dubbing-Clone-V1 video dubbing workflow work?
An application submits a video URL with source and target language parameters, receives a task ID, waits for asynchronous processing, and polls the task status until the translated video is ready.


Sources 2
Provider

About Tencent AI