WAND ASR

WAND-ASR-v1

by Tencent AI · Current and available through Tencent Cloud TokenHub

Tencent Cloud WAND-ASR-v1 is a TokenHub speech-recognition model for audio and video transcription. It accepts URL or Base64 input, supports automatic or specified language selection, and returns text with duration, sentence timestamps, and an SRT subtitle URL. Its reference price is 10 CNY per million tokens, equivalent to approximately 0.0005 CNY per second at the documented usage rate.

Text
WAND-ASR-v1 is Tencent Cloud’s dedicated speech-recognition model for turning audio and video into written text. Its TokenHub interface is designed for synchronous transcription requests and supports URL or Base64 audio input, automatic or specified language selection, sentence timestamps, and SRT subtitle generation. It is a focused ASR tool rather than a general conversational, reasoning, coding, or content-generation model.
Outputs

What WAND-ASR-v1 can produce

Text
Inputs

What it can understand

Audio Video Multimodal input
Model profile

Performance characteristics

7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family WAND ASR
Model type Other
Status Current and available through Tencent Cloud TokenHub
Knowledge cutoff notes

Knowledge-cutoff information is not applicable or publicly specified for this transcription model.

Model notes

The canonical API model identifier is wand-asr-v1. Tencent documents synchronous transcription through /v1/wand/asrproxy/sync_transcribe and also documents an asynchronous ASR workflow using the same model identifier for longer audio files. Inputs can be supplied as an audio or video URL or Base64 audio. The response includes recognized language, duration, full text, sentence timestamps, and an SRT subtitle URL. The speed and cost scores are editorial estimates for speech-recognition use cases, not provider benchmarks. Tencent's documentation does not publish a model-specific release date, knowledge cutoff, context window, maximum output-token limit, fine-tuning support, or structured-output specification.

Cost

Model pricing

Input 10 CNY per 1 million tokens; reference usage is 50 tokens per second, approximately 0.0005 CNY per second
Model guide

WAND-ASR-v1: Tencent’s Speech-to-Text Model for Transcription and Subtitles

WAND-ASR-v1 is Tencent Cloud’s automatic speech-recognition model for converting short audio and video files into text. Available through TokenHub, it accepts a media URL or Base64-encoded audio, can detect the spoken language, and returns a transcript with duration, sentence-level timestamps, and an optional SRT subtitle file.

What is WAND-ASR-v1?

WAND-ASR-v1 is an automatic speech-recognition (ASR) model from Tencent Cloud. ASR systems listen to spoken audio and produce a text transcript. In this case, the model can process audio or video content and return the recognized speech as text through Tencent Cloud’s TokenHub platform.

The model’s canonical API identifier is wand-asr-v1. It is listed in Tencent Cloud’s current TokenHub model catalog and is available through both synchronous transcription and a separately documented asynchronous workflow. The synchronous interface is intended for direct transcription requests, while the asynchronous path is relevant when longer recordings need a job-based processing workflow.

WAND-ASR-v1 should be understood as a specialized transcription model. It is not presented as a general-purpose language model for open-ended conversation, document writing, code generation, or multi-step reasoning.

What the model accepts and returns

A request can provide media through a URL or submit audio as Base64-encoded data. The documented synchronous workflow supports audio and video inputs. When the request does not specify a language, the service can detect the language automatically. The API also supports explicit language selection, including Chinese and English.

The response is useful for more than a plain block of text. According to Tencent Cloud’s documentation, WAND-ASR-v1 can return:

  • The complete recognized transcript
  • The detected or specified language
  • The total duration of the media
  • Sentence-level start and end timestamps
  • An SRT subtitle download URL

Sentence timestamps associate portions of the transcript with their position in the recording. This makes the output more practical for subtitle creation, media review, searchable archives, and applications that need to jump to a particular spoken sentence.

Where WAND-ASR-v1 fits in Tencent Cloud’s lineup

WAND-ASR-v1 occupies the speech-recognition part of Tencent Cloud TokenHub rather than the general text-generation part of the platform. Its role is to convert recorded speech into structured transcription results. The same model identifier appears in Tencent’s synchronous and asynchronous ASR documentation, giving users more than one processing pattern depending on the recording and application workflow.

The supplied documentation does not identify a model-specific release date, knowledge cutoff, context window, or maximum output-token limit. Those omissions are important: the model’s practical boundaries should not be inferred from specifications published for unrelated Tencent models. For long recordings, Tencent’s asynchronous ASR workflow may be more appropriate than a direct synchronous request, but the supplied research does not state a precise maximum recording length.

Pricing and usage economics

Tencent Cloud’s TokenHub pricing lists WAND-ASR-v1 at 10 CNY per 1 million tokens. Tencent’s published usage conversion estimates approximately 50 tokens per second of audio, which corresponds to about 0.0005 CNY per second as a reference calculation.

That conversion implies an approximate cost of 0.03 CNY per minute or 1.80 CNY per hour, but these are arithmetic estimates based on the documented token rate and reference usage, not a separate flat per-minute tariff. Actual charges depend on measured usage and the applicable TokenHub billing rules. Teams evaluating cost should therefore treat the per-second figure as an estimate rather than a guaranteed price for every recording.

The cost profile is most attractive when the task is straightforward transcription and the application does not need a general-purpose model to interpret, summarize, or transform the transcript. If a workflow requires transcription followed by analysis or content generation, those later operations may require additional services or models.

Main strengths

  • Focused speech recognition: The model is designed specifically for converting spoken audio into text rather than dividing its resources among unrelated language-model tasks.
  • Audio and video handling: It can be used with both audio and video content, which suits recorded meetings, interviews, lectures, support calls, and media files.
  • Flexible input delivery: Applications can provide a media URL or Base64-encoded audio.
  • Automatic language detection: The service can select the language when none is supplied, while also allowing an explicit language choice such as Chinese or English.
  • Time-aligned results: Sentence-level timestamps are more useful than an undifferentiated transcript for editing, review, search, and subtitle workflows.
  • Subtitle-ready output: The SRT subtitle URL reduces the amount of post-processing needed for video-captioning applications.

These strengths make WAND-ASR-v1 a practical fit when the desired result is a transcript or subtitle file, not a conversational answer.

Limitations and unsupported or unspecified features

WAND-ASR-v1 has a deliberately narrow role. It produces text from speech, but the supplied research does not support treating it as a reasoning model, coding model, or general-purpose assistant. It does not provide native speech synthesis, music generation, image generation, or video generation.

The model’s documented output is text, timestamps, and subtitle-related data. Although the input may be audio or video, its output is not an audio or video response. This distinction matters when selecting a model for a voice assistant or media-generation pipeline: WAND-ASR-v1 can supply the transcription stage, but it does not itself speak a response or create new media.

Tencent’s available documentation does not publish a model-specific context window, maximum output-token limit, knowledge cutoff, fine-tuning option, structured-output specification, or general-purpose tool-calling capability. The model is also not documented as supporting streaming transcription in the supplied research. These details should be verified with Tencent if they are requirements for a production design.

The synchronous interface may not be the best choice for every recording. Tencent documents an asynchronous ASR workflow using the same model identifier for longer audio files. This gives developers an alternative processing pattern, but the supplied material does not define a precise duration threshold at which asynchronous processing becomes mandatory.

Speed, cost, and capability trade-offs

For a transcription-only job, a dedicated ASR model can be a better fit than a general-purpose multimodal model because the task is narrowly defined and the pricing is expressed around speech-recognition usage. WAND-ASR-v1 also provides transcription-specific outputs such as sentence timestamps and SRT subtitles, which may require additional extraction or post-processing from a general model.

Its trade-off is breadth. A general-purpose model may be more suitable when the same request must interpret images, reason over several documents, write code, summarize a transcript, or carry out a tool-based action. WAND-ASR-v1 should instead be viewed as one component in a pipeline: it transcribes the recording, after which another service can summarize, classify, translate, or search the resulting text.

The research assigns WAND-ASR-v1 an editorial speed score of 7 and cost score of 8. These are internal comparative estimates for speech-recognition use cases, not Tencent benchmarks or provider-published ratings. They indicate a favorable practical balance for focused transcription, but they should not be used as measured latency or guaranteed cost commitments.

Best use cases

WAND-ASR-v1 is well suited to applications that need a text representation of recorded speech together with timing information. Suitable examples include:

  • Transcribing meetings, interviews, and training sessions
  • Creating subtitles for recorded video
  • Indexing media libraries so users can search spoken content
  • Generating transcripts of recorded customer-service conversations
  • Processing short audio or video files in a synchronous workflow
  • Preparing time-aligned text for media editing and review

For longer recordings, an asynchronous workflow using the same model identifier may provide a better operational fit than waiting for a synchronous response. The right choice depends on the recording length, application timeout requirements, and Tencent Cloud’s current service limits.

When should you choose WAND-ASR-v1?

Choose WAND-ASR-v1 when the central requirement is reliable access to a transcript, detected language, sentence timestamps, or SRT subtitles through Tencent Cloud TokenHub. It is especially appropriate when audio or video is already available at a URL or can be encoded as Base64 audio, and when the application benefits from a focused ASR endpoint rather than a broad conversational model.

Another option may be more appropriate when the application needs real-time streaming, a published maximum input or output limit, fine-tuning, structured outputs, function calling, or a model that can directly generate speech. The supplied research does not verify those capabilities for WAND-ASR-v1. Similarly, use a broader language or multimodal system after transcription when the main task is analysis, reasoning, coding, summarization, or content generation rather than speech recognition itself.

Bottom line

WAND-ASR-v1 is a focused Tencent Cloud TokenHub model for converting audio and video into usable text. Its most distinctive practical features are automatic or specified language handling, sentence-level timestamps, and SRT subtitle output. At a reference rate of 10 CNY per million tokens and approximately 50 tokens per second, it offers a clear usage-based pricing model for transcription workloads.

Its value comes from specialization, not from general-purpose intelligence. It is a sensible choice for transcription and subtitle pipelines, while workflows requiring conversation, reasoning, coding, tool use, or native media generation should add or select other services.


Answers to Frequently Asked Questions

What are the main limitations of WAND-ASR-v1?
WAND-ASR-v1 is specialized for speech recognition and is not a general-purpose conversational, reasoning, coding, or media-generation model. The supplied documentation does not verify streaming transcription, a model-specific context window, maximum output-token limit, fine-tuning, structured outputs, or tool calling. Longer recordings may be better suited to Tencent’s asynchronous ASR workflow.
Does WAND-ASR-v1 support automatic language detection and subtitles?
Yes. WAND-ASR-v1 can detect the language automatically when no language is specified, and it supports explicit language selection such as Chinese and English. Its sentence-level timestamps and SRT subtitle URL make it suitable for captioning and time-aligned media workflows.
How much does WAND-ASR-v1 cost?
Tencent Cloud TokenHub lists WAND-ASR-v1 at 10 CNY per 1 million tokens. Using Tencent’s estimate of approximately 50 tokens per second, the reference cost is about 0.0005 CNY per second, 0.03 CNY per minute, or 1.80 CNY per hour. Actual charges depend on measured usage and TokenHub billing rules.
What is WAND-ASR-v1 used for?
WAND-ASR-v1 is Tencent Cloud’s automatic speech-recognition model for converting audio or video into text. It is designed for transcription, subtitle creation, searchable media archives, meeting records, interviews, lectures, and recorded customer-service conversations.
What inputs and outputs does WAND-ASR-v1 support?
WAND-ASR-v1 accepts audio or video through a URL, and audio can also be submitted as Base64-encoded data. It can return a complete transcript, detected or specified language, media duration, sentence-level start and end timestamps, and an SRT subtitle download URL.


Sources 4
Provider

About Tencent AI