What is WAND-Dubbing-Clone-v2?
WAND-Dubbing-Clone-v2 is Tencent Cloud’s asynchronous AI model for translating and dubbing video. A user supplies a publicly accessible video URL or a base64-encoded video payload, optionally specifies the source and destination languages, and submits a dubbing task. The service processes that task in the background and returns a translated video URL when it is complete.
The model combines several steps that are often handled by separate tools: translating the dialogue, generating speech in the target language, preserving important characteristics of the original speaker’s voice, and producing a synchronized dubbed video with subtitles. This makes it a localization product rather than a general-purpose text or speech model.
The canonical model identifier is wand-dubbing-clone-v2. Tencent Cloud lists it in TokenHub’s current voice-model catalog and exposes it through an asynchronous workflow rather than a streaming interaction.
What the model does in practice
WAND-Dubbing-Clone-v2 is intended to turn a video in one language into a version that can be understood by viewers who speak another. For example, an organization could use it to localize an online course, a product explainer, a short social video, or other media without recording every target-language version manually.
The model translates both the video’s speech and subtitles, generates the target-language dubbing, and returns a video that includes the translated result. Tencent Cloud describes the v2 model as improving the naturalness of translated phrasing, the preservation of the original speaker’s emotional and prosodic delivery, and stability when processing longer videos.
“Prosody” refers to delivery features such as rhythm, stress, intonation, and pacing. Preserving these characteristics can make a dubbed speaker sound more like the source performance than a basic translation-and-text-to-speech pipeline would. These are provider-described capabilities; the supplied research does not include an independent benchmark quantifying the improvement.
Inputs, outputs, and asynchronous workflow
The primary input is video. The documented input options are:
- A publicly accessible video URL.
- A base64-encoded video payload up to 100 MB.
- Optional source-language and destination-language codes.
The output is a translated and dubbed video URL, with subtitles included as part of the service’s video-localization result. The model also produces target-language speech as part of that video output. It is therefore best understood as a video-to-video service with generated audio, not as a conventional text-in/text-out model.
Processing is asynchronous. A client submits a task through Tencent Cloud’s /v1/wand/dubbing/async_generation endpoint and then polls a task endpoint for completion. This workflow is appropriate for production jobs that may take time to process, but it is not designed for live conversation, instant speech-to-speech exchange, or token-by-token streaming. The supplied research does not specify a maximum video duration, a completion-time guarantee, or a separate output-size limit.
Supported languages
Tencent Cloud lists support for a broad set of languages and language variants, including Chinese, English, Japanese, German, French, Korean, Russian, Ukrainian, Portuguese, Italian, Spanish, Indonesian, Dutch, Turkish, Filipino, Malay, Greek, Finnish, Croatian, Slovak, Polish, Swedish, Hindi, Bulgarian, Romanian, Arabic, Czech, Danish, Tamil, Hungarian, Vietnamese, Thai, and Cantonese.
Language availability should still be checked against the current Tencent Cloud documentation before a production rollout. A language appearing in the catalog does not by itself establish that every language pair has identical quality, voice behavior, or subtitle performance.
Main strengths
Voice and emotional delivery preservation
The most distinctive feature is the attempt to preserve the source speaker’s voice characteristics, emotion, tone, and delivery while rendering the dialogue in another language. This is more specific than simply translating subtitles or generating a generic target-language voice. It is especially relevant when the identity and performance of the speaker contribute to the value of the video.
End-to-end video localization
Because the service handles translation, speech generation, and dubbed-video output in one task, it can reduce the need to manually connect separate transcription, translation, text-to-speech, and media-editing systems. That does not eliminate the need for review, but it can simplify the path from source video to localized version.
Improved stability for longer videos
Tencent Cloud positions the v2 release as more stable for longer videos. This matters for courses, lectures, interviews, and other content that is too long for a short-form-only workflow. The research does not provide a supported maximum duration or a measured failure rate, so “longer” should be treated as a provider claim rather than a guaranteed technical limit.
Broad language coverage
The listed language set covers major European and Asian languages as well as Arabic and several other regional languages. This makes the model a candidate for organizations that need to localize the same video into multiple markets rather than only one common language pair.
Limitations and trade-offs
The model’s specialization is also its main limitation. It does not appear to be intended for general conversation, reasoning over documents, coding, image generation, transcription-only work, or standalone text-to-speech. It accepts video as the central input and is not documented here as accepting independent audio input or text prompts as a replacement for video.
The asynchronous design adds operational delay. A live application cannot treat this service like a streaming speech API, because the normal pattern is to submit a task and poll for its status. Applications also need to manage video hosting or base64 encoding, task state, retries, and the returned media URL.
There are no supplied values for context length or maximum output tokens. Those fields are not directly applicable in the same way they are for a text-generation model, but users should not infer unlimited video length or unlimited output size. The documented base64 input limit is up to 100 MB; the research does not state whether the same limit applies to URL-based inputs.
Automatic dubbing should also be reviewed before publication. Translation nuance, names, specialized terminology, timing, subtitle segmentation, and emotional delivery can vary by language and content type. Human review remains sensible for education, legal or regulated material, advertising, and media where voice identity or meaning is especially sensitive.
Pricing
Tencent Cloud’s listed pricing is based on the video dubbing task and resolution rather than a conventional separate input-token and output-token price:
| Video resolution | Listed price |
|---|---|
| 720p or lower | ¥0.2061 per second |
| 1080p | ¥0.2311 per second |
| 2K or higher | ¥0.2811 per second |
TokenHub also lists an underlying rate of ¥10 per million tokens, but the supplied pricing information says that WAND-Dubbing-Clone-v2 is charged according to the resolution-based dubbing-task rates above. For budgeting, the practical calculation is therefore video duration multiplied by the applicable per-second resolution price, subject to Tencent Cloud’s current billing terms.
For example, a 10-minute video at 1080p would correspond to 600 seconds multiplied by ¥0.2311, or ¥138.66 before any applicable account-level terms, taxes, discounts, or other charges. This is an arithmetic illustration of the listed rate, not a guarantee of the final invoice.
Capabilities at a glance
- Input: Video URL or base64-encoded video payload.
- Input limit documented in the supplied research: Up to 100 MB for base64 video.
- Output: Translated dubbed video with subtitles and generated target-language speech.
- Processing: Asynchronous task submission and polling.
- Streaming: Not supported in the supplied specifications.
- Tool or function calling: Not supported as a model capability.
- Structured JSON output: Not documented; the service returns a media result rather than a text-generation schema.
- Fine-tuning: Not listed in the supplied specifications.
- Web search: Not a model capability.
The model’s modality profile is narrow but concrete: video is the key input, while video and generated speech are part of the output. It should not be evaluated using the same criteria as a multimodal reasoning model that accepts images, text, and audio for open-ended analysis.
Reasoning, coding, and speed considerations
WAND-Dubbing-Clone-v2 is not a reasoning model in the usual language-model sense. Its job is to translate and render video dialogue, not to solve multi-step analytical problems. The supplied editorial assessment assigns it a reasoning score of 2 out of 10 and a coding score of 1 out of 10. These are internal comparative evaluations, not Tencent Cloud benchmarks or provider-published quality ratings.
The editorial speed assessment is 6 out of 10, reflecting a useful but asynchronous workflow rather than real-time interaction. The model is likely a better fit when completed localized media matters more than immediate response. The editorial cost assessment is 7 out of 10 because the listed per-second rates can be attractive for automated localization, particularly at lower resolutions, but total cost rises with duration and resolution.
These scores should be used only as directional trade-off indicators. The verified facts are the asynchronous design, resolution-based rates, supported input format, and documented dubbing behavior.
When to choose WAND-Dubbing-Clone-v2
Choose this model when the main deliverable is a translated video rather than a transcript, translation string, isolated audio file, or conversational response. It is particularly suitable for:
- Localizing online courses and instructional videos.
- Creating multilingual versions of short-form video.
- Translating product demonstrations and explainers.
- Preparing media content for several language markets.
- Preserving a presenter’s recognizable voice and emotional delivery.
- Processing longer videos through an asynchronous production pipeline.
Another option may be more appropriate when the requirement is real-time speech-to-speech interaction, transcription only, generic text-to-speech, unrestricted text reasoning, software development, or image generation. A separate translation service may also be preferable when the team needs fine-grained control over terminology and wants to arrange its own voice and video-editing pipeline.
Bottom line
WAND-Dubbing-Clone-v2 is a focused Tencent Cloud model for multilingual video localization. Its value comes from combining translation, voice-preserving speech generation, subtitles, and dubbed-video delivery in an asynchronous task. The clearest reasons to select it are the broad listed language coverage, emphasis on preserving speaker delivery, support for longer-video workflows, and transparent resolution-based pricing. The clearest reasons to look elsewhere are the lack of real-time processing, the narrow video-centered scope, and the absence of documented limits for maximum duration or token-style output controls.

