WAND-Dubbing-Clone

WAND-Dubbing-Clone-v2

by Tencent AI · Current and available through Tencent Cloud TokenHub as an asynchronous AI dubbing model

Tencent Cloud’s WAND-Dubbing-Clone-v2 translates video dialogue, generates target-language speech that preserves the original speaker’s voice and delivery, and returns a dubbed video with subtitles. It supports video URLs or base64 uploads up to 100 MB, uses asynchronous task processing, and charges by video resolution and duration.

Video generation Speech Reasoning Coding
WAND-Dubbing-Clone-v2 is Tencent Cloud WAND’s upgraded model for translating and dubbing video. Rather than producing text or speech as separate outputs, it processes a video task from submission through translated dialogue, voice generation, and dubbed-video delivery. The model is designed for asynchronous use, supports a broad list of source and target languages, and improves phrasing, emotional delivery, and stability for longer videos. Its main trade-off is specialization: it is useful for video localization, but it is not a general conversational model, standalone text-to-speech system, or real-time speech service.
Outputs

What WAND-Dubbing-Clone-v2 can produce

Video generation Speech
Inputs

What it can understand

Video Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
6/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family WAND-Dubbing-Clone
Model type Other
Status Current and available through Tencent Cloud TokenHub as an asynchronous AI dubbing model
Model notes

Canonical model identifier: wand-dubbing-clone-v2. The model accepts a publicly accessible video URL or a base64-encoded video payload up to 100 MB, along with optional source and destination language codes. It runs through the asynchronous /v1/wand/dubbing/async_generation endpoint; clients submit a task and poll a task endpoint for completion. The service translates video subtitles and speech, generates target-language dubbing, preserves the source speaker's emotional and prosodic characteristics, and returns a translated video URL. Supported languages listed by Tencent Cloud include Chinese, English, Japanese, German, French, Korean, Russian, Ukrainian, Portuguese, Italian, Spanish, Indonesian, Dutch, Turkish, Filipino, Malay, Greek, Finnish, Croatian, Slovak, Polish, Swedish, Hindi, Bulgarian, Romanian, Arabic, Czech, Danish, Tamil, Hungarian, Vietnamese, Thai, and Cantonese. Pricing is resolution-dependent and applies to the dubbing task rather than separate conventional input and output token prices.

Cost

Model pricing

Input ¥0.2061 per second for 720p or lower; ¥0.2311 per second at 1080p; ¥0.2811 per second at 2K or higher. TokenHub lists the underlying rate as ¥10 per million tokens.
Model guide

WAND-Dubbing-Clone-v2: Tencent Cloud’s Video Translation and Voice-Preserving Dubbing Model

Tencent Cloud’s WAND-Dubbing-Clone-v2 is an asynchronous video localization model. It translates spoken dialogue, recreates the result in a target language with closely matched voice characteristics and delivery, and returns a dubbed video with subtitles. Its strongest use cases are multilingual courses, online video, short-form content, and longer media localization projects where preserving the original speaker’s emotion and prosody matters.

What is WAND-Dubbing-Clone-v2?

WAND-Dubbing-Clone-v2 is Tencent Cloud’s asynchronous AI model for translating and dubbing video. A user supplies a publicly accessible video URL or a base64-encoded video payload, optionally specifies the source and destination languages, and submits a dubbing task. The service processes that task in the background and returns a translated video URL when it is complete.

The model combines several steps that are often handled by separate tools: translating the dialogue, generating speech in the target language, preserving important characteristics of the original speaker’s voice, and producing a synchronized dubbed video with subtitles. This makes it a localization product rather than a general-purpose text or speech model.

The canonical model identifier is wand-dubbing-clone-v2. Tencent Cloud lists it in TokenHub’s current voice-model catalog and exposes it through an asynchronous workflow rather than a streaming interaction.

What the model does in practice

WAND-Dubbing-Clone-v2 is intended to turn a video in one language into a version that can be understood by viewers who speak another. For example, an organization could use it to localize an online course, a product explainer, a short social video, or other media without recording every target-language version manually.

The model translates both the video’s speech and subtitles, generates the target-language dubbing, and returns a video that includes the translated result. Tencent Cloud describes the v2 model as improving the naturalness of translated phrasing, the preservation of the original speaker’s emotional and prosodic delivery, and stability when processing longer videos.

“Prosody” refers to delivery features such as rhythm, stress, intonation, and pacing. Preserving these characteristics can make a dubbed speaker sound more like the source performance than a basic translation-and-text-to-speech pipeline would. These are provider-described capabilities; the supplied research does not include an independent benchmark quantifying the improvement.

Inputs, outputs, and asynchronous workflow

The primary input is video. The documented input options are:

  • A publicly accessible video URL.
  • A base64-encoded video payload up to 100 MB.
  • Optional source-language and destination-language codes.

The output is a translated and dubbed video URL, with subtitles included as part of the service’s video-localization result. The model also produces target-language speech as part of that video output. It is therefore best understood as a video-to-video service with generated audio, not as a conventional text-in/text-out model.

Processing is asynchronous. A client submits a task through Tencent Cloud’s /v1/wand/dubbing/async_generation endpoint and then polls a task endpoint for completion. This workflow is appropriate for production jobs that may take time to process, but it is not designed for live conversation, instant speech-to-speech exchange, or token-by-token streaming. The supplied research does not specify a maximum video duration, a completion-time guarantee, or a separate output-size limit.

Supported languages

Tencent Cloud lists support for a broad set of languages and language variants, including Chinese, English, Japanese, German, French, Korean, Russian, Ukrainian, Portuguese, Italian, Spanish, Indonesian, Dutch, Turkish, Filipino, Malay, Greek, Finnish, Croatian, Slovak, Polish, Swedish, Hindi, Bulgarian, Romanian, Arabic, Czech, Danish, Tamil, Hungarian, Vietnamese, Thai, and Cantonese.

Language availability should still be checked against the current Tencent Cloud documentation before a production rollout. A language appearing in the catalog does not by itself establish that every language pair has identical quality, voice behavior, or subtitle performance.

Main strengths

Voice and emotional delivery preservation

The most distinctive feature is the attempt to preserve the source speaker’s voice characteristics, emotion, tone, and delivery while rendering the dialogue in another language. This is more specific than simply translating subtitles or generating a generic target-language voice. It is especially relevant when the identity and performance of the speaker contribute to the value of the video.

End-to-end video localization

Because the service handles translation, speech generation, and dubbed-video output in one task, it can reduce the need to manually connect separate transcription, translation, text-to-speech, and media-editing systems. That does not eliminate the need for review, but it can simplify the path from source video to localized version.

Improved stability for longer videos

Tencent Cloud positions the v2 release as more stable for longer videos. This matters for courses, lectures, interviews, and other content that is too long for a short-form-only workflow. The research does not provide a supported maximum duration or a measured failure rate, so “longer” should be treated as a provider claim rather than a guaranteed technical limit.

Broad language coverage

The listed language set covers major European and Asian languages as well as Arabic and several other regional languages. This makes the model a candidate for organizations that need to localize the same video into multiple markets rather than only one common language pair.

Limitations and trade-offs

The model’s specialization is also its main limitation. It does not appear to be intended for general conversation, reasoning over documents, coding, image generation, transcription-only work, or standalone text-to-speech. It accepts video as the central input and is not documented here as accepting independent audio input or text prompts as a replacement for video.

The asynchronous design adds operational delay. A live application cannot treat this service like a streaming speech API, because the normal pattern is to submit a task and poll for its status. Applications also need to manage video hosting or base64 encoding, task state, retries, and the returned media URL.

There are no supplied values for context length or maximum output tokens. Those fields are not directly applicable in the same way they are for a text-generation model, but users should not infer unlimited video length or unlimited output size. The documented base64 input limit is up to 100 MB; the research does not state whether the same limit applies to URL-based inputs.

Automatic dubbing should also be reviewed before publication. Translation nuance, names, specialized terminology, timing, subtitle segmentation, and emotional delivery can vary by language and content type. Human review remains sensible for education, legal or regulated material, advertising, and media where voice identity or meaning is especially sensitive.

Pricing

Tencent Cloud’s listed pricing is based on the video dubbing task and resolution rather than a conventional separate input-token and output-token price:

Video resolutionListed price
720p or lower¥0.2061 per second
1080p¥0.2311 per second
2K or higher¥0.2811 per second

TokenHub also lists an underlying rate of ¥10 per million tokens, but the supplied pricing information says that WAND-Dubbing-Clone-v2 is charged according to the resolution-based dubbing-task rates above. For budgeting, the practical calculation is therefore video duration multiplied by the applicable per-second resolution price, subject to Tencent Cloud’s current billing terms.

For example, a 10-minute video at 1080p would correspond to 600 seconds multiplied by ¥0.2311, or ¥138.66 before any applicable account-level terms, taxes, discounts, or other charges. This is an arithmetic illustration of the listed rate, not a guarantee of the final invoice.

Capabilities at a glance

  • Input: Video URL or base64-encoded video payload.
  • Input limit documented in the supplied research: Up to 100 MB for base64 video.
  • Output: Translated dubbed video with subtitles and generated target-language speech.
  • Processing: Asynchronous task submission and polling.
  • Streaming: Not supported in the supplied specifications.
  • Tool or function calling: Not supported as a model capability.
  • Structured JSON output: Not documented; the service returns a media result rather than a text-generation schema.
  • Fine-tuning: Not listed in the supplied specifications.
  • Web search: Not a model capability.

The model’s modality profile is narrow but concrete: video is the key input, while video and generated speech are part of the output. It should not be evaluated using the same criteria as a multimodal reasoning model that accepts images, text, and audio for open-ended analysis.

Reasoning, coding, and speed considerations

WAND-Dubbing-Clone-v2 is not a reasoning model in the usual language-model sense. Its job is to translate and render video dialogue, not to solve multi-step analytical problems. The supplied editorial assessment assigns it a reasoning score of 2 out of 10 and a coding score of 1 out of 10. These are internal comparative evaluations, not Tencent Cloud benchmarks or provider-published quality ratings.

The editorial speed assessment is 6 out of 10, reflecting a useful but asynchronous workflow rather than real-time interaction. The model is likely a better fit when completed localized media matters more than immediate response. The editorial cost assessment is 7 out of 10 because the listed per-second rates can be attractive for automated localization, particularly at lower resolutions, but total cost rises with duration and resolution.

These scores should be used only as directional trade-off indicators. The verified facts are the asynchronous design, resolution-based rates, supported input format, and documented dubbing behavior.

When to choose WAND-Dubbing-Clone-v2

Choose this model when the main deliverable is a translated video rather than a transcript, translation string, isolated audio file, or conversational response. It is particularly suitable for:

  • Localizing online courses and instructional videos.
  • Creating multilingual versions of short-form video.
  • Translating product demonstrations and explainers.
  • Preparing media content for several language markets.
  • Preserving a presenter’s recognizable voice and emotional delivery.
  • Processing longer videos through an asynchronous production pipeline.

Another option may be more appropriate when the requirement is real-time speech-to-speech interaction, transcription only, generic text-to-speech, unrestricted text reasoning, software development, or image generation. A separate translation service may also be preferable when the team needs fine-grained control over terminology and wants to arrange its own voice and video-editing pipeline.

Bottom line

WAND-Dubbing-Clone-v2 is a focused Tencent Cloud model for multilingual video localization. Its value comes from combining translation, voice-preserving speech generation, subtitles, and dubbed-video delivery in an asynchronous task. The clearest reasons to select it are the broad listed language coverage, emphasis on preserving speaker delivery, support for longer-video workflows, and transparent resolution-based pricing. The clearest reasons to look elsewhere are the lack of real-time processing, the narrow video-centered scope, and the absence of documented limits for maximum duration or token-style output controls.


Answers to Frequently Asked Questions

What are the main limitations of WAND-Dubbing-Clone-v2?
WAND-Dubbing-Clone-v2 is designed for asynchronous video localization rather than live conversation, streaming speech-to-speech interaction, transcription-only tasks, general reasoning, coding, or standalone text-to-speech. The supplied specifications document a 100 MB limit for base64 video but do not state a maximum video duration, completion-time guarantee, or separate output-size limit, so human review is recommended for sensitive or specialized content.
How much does WAND-Dubbing-Clone-v2 cost?
Tencent Cloud lists resolution-based pricing: ¥0.2061 per second for 720p or lower, ¥0.2311 per second for 1080p, and ¥0.2811 per second for 2K or higher. For example, a 10-minute 1080p video corresponds to 600 × ¥0.2311, or ¥138.66 before applicable taxes, discounts, account terms, or other charges.
What inputs and outputs does WAND-Dubbing-Clone-v2 support?
The model accepts a public video URL or a base64-encoded video payload of up to 100 MB, along with optional language codes. It produces a translated and dubbed video URL that includes subtitles and generated speech in the target language.
What is WAND-Dubbing-Clone-v2 used for?
WAND-Dubbing-Clone-v2 is Tencent Cloud’s asynchronous AI model for translating and dubbing videos. It translates dialogue and subtitles, generates target-language speech, preserves important characteristics of the original speaker’s voice and delivery, and returns a localized dubbed video.
How does the WAND-Dubbing-Clone-v2 workflow work?
Users submit a publicly accessible video URL or a base64-encoded video payload, optionally provide source and destination language codes, and create a task through Tencent Cloud’s /v1/wand/dubbing/async_generation endpoint. The client then polls a task endpoint until the translated video URL is available.


Sources 4
Provider

About Tencent AI