What is WAND-ASR-v1?
WAND-ASR-v1 is an automatic speech-recognition (ASR) model from Tencent Cloud. ASR systems listen to spoken audio and produce a text transcript. In this case, the model can process audio or video content and return the recognized speech as text through Tencent Cloud’s TokenHub platform.
The model’s canonical API identifier is wand-asr-v1. It is listed in Tencent Cloud’s current TokenHub model catalog and is available through both synchronous transcription and a separately documented asynchronous workflow. The synchronous interface is intended for direct transcription requests, while the asynchronous path is relevant when longer recordings need a job-based processing workflow.
WAND-ASR-v1 should be understood as a specialized transcription model. It is not presented as a general-purpose language model for open-ended conversation, document writing, code generation, or multi-step reasoning.
What the model accepts and returns
A request can provide media through a URL or submit audio as Base64-encoded data. The documented synchronous workflow supports audio and video inputs. When the request does not specify a language, the service can detect the language automatically. The API also supports explicit language selection, including Chinese and English.
The response is useful for more than a plain block of text. According to Tencent Cloud’s documentation, WAND-ASR-v1 can return:
- The complete recognized transcript
- The detected or specified language
- The total duration of the media
- Sentence-level start and end timestamps
- An SRT subtitle download URL
Sentence timestamps associate portions of the transcript with their position in the recording. This makes the output more practical for subtitle creation, media review, searchable archives, and applications that need to jump to a particular spoken sentence.
Where WAND-ASR-v1 fits in Tencent Cloud’s lineup
WAND-ASR-v1 occupies the speech-recognition part of Tencent Cloud TokenHub rather than the general text-generation part of the platform. Its role is to convert recorded speech into structured transcription results. The same model identifier appears in Tencent’s synchronous and asynchronous ASR documentation, giving users more than one processing pattern depending on the recording and application workflow.
The supplied documentation does not identify a model-specific release date, knowledge cutoff, context window, or maximum output-token limit. Those omissions are important: the model’s practical boundaries should not be inferred from specifications published for unrelated Tencent models. For long recordings, Tencent’s asynchronous ASR workflow may be more appropriate than a direct synchronous request, but the supplied research does not state a precise maximum recording length.
Pricing and usage economics
Tencent Cloud’s TokenHub pricing lists WAND-ASR-v1 at 10 CNY per 1 million tokens. Tencent’s published usage conversion estimates approximately 50 tokens per second of audio, which corresponds to about 0.0005 CNY per second as a reference calculation.
That conversion implies an approximate cost of 0.03 CNY per minute or 1.80 CNY per hour, but these are arithmetic estimates based on the documented token rate and reference usage, not a separate flat per-minute tariff. Actual charges depend on measured usage and the applicable TokenHub billing rules. Teams evaluating cost should therefore treat the per-second figure as an estimate rather than a guaranteed price for every recording.
The cost profile is most attractive when the task is straightforward transcription and the application does not need a general-purpose model to interpret, summarize, or transform the transcript. If a workflow requires transcription followed by analysis or content generation, those later operations may require additional services or models.
Main strengths
- Focused speech recognition: The model is designed specifically for converting spoken audio into text rather than dividing its resources among unrelated language-model tasks.
- Audio and video handling: It can be used with both audio and video content, which suits recorded meetings, interviews, lectures, support calls, and media files.
- Flexible input delivery: Applications can provide a media URL or Base64-encoded audio.
- Automatic language detection: The service can select the language when none is supplied, while also allowing an explicit language choice such as Chinese or English.
- Time-aligned results: Sentence-level timestamps are more useful than an undifferentiated transcript for editing, review, search, and subtitle workflows.
- Subtitle-ready output: The SRT subtitle URL reduces the amount of post-processing needed for video-captioning applications.
These strengths make WAND-ASR-v1 a practical fit when the desired result is a transcript or subtitle file, not a conversational answer.
Limitations and unsupported or unspecified features
WAND-ASR-v1 has a deliberately narrow role. It produces text from speech, but the supplied research does not support treating it as a reasoning model, coding model, or general-purpose assistant. It does not provide native speech synthesis, music generation, image generation, or video generation.
The model’s documented output is text, timestamps, and subtitle-related data. Although the input may be audio or video, its output is not an audio or video response. This distinction matters when selecting a model for a voice assistant or media-generation pipeline: WAND-ASR-v1 can supply the transcription stage, but it does not itself speak a response or create new media.
Tencent’s available documentation does not publish a model-specific context window, maximum output-token limit, knowledge cutoff, fine-tuning option, structured-output specification, or general-purpose tool-calling capability. The model is also not documented as supporting streaming transcription in the supplied research. These details should be verified with Tencent if they are requirements for a production design.
The synchronous interface may not be the best choice for every recording. Tencent documents an asynchronous ASR workflow using the same model identifier for longer audio files. This gives developers an alternative processing pattern, but the supplied material does not define a precise duration threshold at which asynchronous processing becomes mandatory.
Speed, cost, and capability trade-offs
For a transcription-only job, a dedicated ASR model can be a better fit than a general-purpose multimodal model because the task is narrowly defined and the pricing is expressed around speech-recognition usage. WAND-ASR-v1 also provides transcription-specific outputs such as sentence timestamps and SRT subtitles, which may require additional extraction or post-processing from a general model.
Its trade-off is breadth. A general-purpose model may be more suitable when the same request must interpret images, reason over several documents, write code, summarize a transcript, or carry out a tool-based action. WAND-ASR-v1 should instead be viewed as one component in a pipeline: it transcribes the recording, after which another service can summarize, classify, translate, or search the resulting text.
The research assigns WAND-ASR-v1 an editorial speed score of 7 and cost score of 8. These are internal comparative estimates for speech-recognition use cases, not Tencent benchmarks or provider-published ratings. They indicate a favorable practical balance for focused transcription, but they should not be used as measured latency or guaranteed cost commitments.
Best use cases
WAND-ASR-v1 is well suited to applications that need a text representation of recorded speech together with timing information. Suitable examples include:
- Transcribing meetings, interviews, and training sessions
- Creating subtitles for recorded video
- Indexing media libraries so users can search spoken content
- Generating transcripts of recorded customer-service conversations
- Processing short audio or video files in a synchronous workflow
- Preparing time-aligned text for media editing and review
For longer recordings, an asynchronous workflow using the same model identifier may provide a better operational fit than waiting for a synchronous response. The right choice depends on the recording length, application timeout requirements, and Tencent Cloud’s current service limits.
When should you choose WAND-ASR-v1?
Choose WAND-ASR-v1 when the central requirement is reliable access to a transcript, detected language, sentence timestamps, or SRT subtitles through Tencent Cloud TokenHub. It is especially appropriate when audio or video is already available at a URL or can be encoded as Base64 audio, and when the application benefits from a focused ASR endpoint rather than a broad conversational model.
Another option may be more appropriate when the application needs real-time streaming, a published maximum input or output limit, fine-tuning, structured outputs, function calling, or a model that can directly generate speech. The supplied research does not verify those capabilities for WAND-ASR-v1. Similarly, use a broader language or multimodal system after transcription when the main task is analysis, reasoning, coding, summarization, or content generation rather than speech recognition itself.
Bottom line
WAND-ASR-v1 is a focused Tencent Cloud TokenHub model for converting audio and video into usable text. Its most distinctive practical features are automatic or specified language handling, sentence-level timestamps, and SRT subtitle output. At a reference rate of 10 CNY per million tokens and approximately 50 tokens per second, it offers a clear usage-based pricing model for transcription workloads.
Its value comes from specialization, not from general-purpose intelligence. It is a sensible choice for transcription and subtitle pipelines, while workflows requiring conversation, reasoning, coding, tool use, or native media generation should add or select other services.

