What Hy-ASR-3.0-Preview is
Hy-ASR-3.0-Preview is Tencent’s preview automatic speech-recognition (ASR) model. ASR systems analyze spoken audio and return a written transcription; this model is intended specifically for that task rather than for open-ended conversation, coding, image generation, or general text completion.
Within Tencent’s current product lineup, it sits in the Hunyuan speech-recognition offering and is exposed through several Tencent Cloud services. The preview and internal-test status means that availability and feature support are more limited than they would be for a mature, unrestricted production model. Tencent’s documentation identifies the real-time model as Hy-ASR-3.0-preview; TokenHub uses the lowercase identifier hy-asr-3.0-preview.
Languages and recognition strengths
The documented language coverage includes Mandarin Chinese, English, and 20 Chinese dialects or regional varieties. These include Cantonese, Northeastern Mandarin, Henan, Shaanxi, Chengdu, Chongqing, Wuhan, Guiyang, Qingdao, Jinan, Changsha, Hefei, Hebei, Kunming, Lanzhou, Yinchuan, Nanchang, Beijing, Sichuan, and Tianjin varieties.
This coverage makes the model more specialized than a speech-to-text service focused only on standard Mandarin. It is also intended to handle Chinese-English mixed speech, which is common in technical discussions, customer calls, livestreams, and workplace conversations. Tencent describes improvements in contextual disambiguation and homophone correction, meaning the model attempts to use surrounding words to choose a more plausible transcription instead of treating every sound in isolation.
Hotword injection is another documented capability. A hotword is a user-supplied term that the recognizer should pay particular attention to, such as a product name, company name, person’s name, medical term, or specialist vocabulary. This can be useful when ordinary language models might replace an uncommon term with a more familiar-sounding word.
Supported inputs and outputs
Hy-ASR-3.0-Preview accepts audio and produces text. It does not provide direct image, video, audio, music, embedding, or general structured-object output. Video can still be processed indirectly through Tencent Cloud media services when the service extracts or transcribes the audio track for subtitles.
The real-time preview interface is documented for 16 kHz, mono PCM audio. The preview real-time service supports audio of up to 60 seconds, or one minute, per recognition session. This is a significant operational constraint: an application handling a long meeting, call, or livestream must divide the recording into suitable segments or use an asynchronous or media-processing workflow where supported.
Output can include a full transcription and sentence timestamps. Some interfaces also provide word-level timestamps, while applicable Tencent Cloud services can produce subtitle files. The supplied documentation does not publish a language-model context window, maximum output-token value, or a single universal output limit across all interfaces. Limits therefore depend on the selected Tencent Cloud service and workflow.
Access methods and practical workflows
For live applications, the model can be selected through Tencent Cloud’s real-time speech-recognition WebSocket API. WebSocket is a persistent connection that allows audio to be sent incrementally and transcription results to be returned as recognition proceeds. This makes the model suitable for live captions, voice-command interfaces, and interactive audio applications.
For files and media operations, Tencent Cloud Media Processing supports Hy ASR 3.0 preview integration for use cases such as subtitle extraction and synchronous audio recognition. TokenHub provides a synchronous transcription endpoint using the lowercase model identifier. These interfaces are useful when the application does not need to display partial results immediately or when audio is already stored as part of a media-processing job.
The exact request format, authentication requirements, regional availability, and service-level limits belong to the individual Tencent Cloud interface rather than to the model alone. Developers should therefore select the service first—real-time recognition, media processing, or TokenHub—and then follow that service’s current documentation.
Preview limitations and unsupported features
The most important limitation is the one-minute real-time session restriction documented for the preview interface. It makes the model a better fit for short live segments or applications that can manage segmentation than for a single unrestricted audio stream.
The internal-test version does not support speaker diarization, voice activity detection (VAD), vocabulary replacement, or noise-threshold parameters. Speaker diarization identifies which person is speaking, so applications that require labels such as “Speaker 1” and “Speaker 2” should not assume that this preview model provides them. VAD detects when speech starts and stops; without the documented VAD feature, the surrounding application or service may need to manage silence handling separately.
Vocabulary replacement and noise-threshold controls are also unavailable in the internal-test version. This limits customization for specialized terminology and for applications that need precise control over how background noise affects recognition. Tencent does describe improved stability in noisy, whispered, and otherwise difficult acoustic conditions, but that provider claim should not be treated as a substitute for a published benchmark or a guarantee for every recording environment.
Reasoning, coding, and tool support
Hy-ASR-3.0-Preview is not a reasoning or coding model in the usual language-model sense. Its contextual understanding is used to improve transcription—for example, resolving homophones or preserving mixed Chinese-English phrases—not to solve multi-step problems or generate software.
The supplied model assessment assigns a low comparative coding score and does not identify tool or function-calling support. These are editorial evaluations and catalog classifications, not Tencent-published benchmarks. The model should be treated as a transcription component that can be called by an application, not as an agent that independently invokes tools, browses the web, writes code, or performs business actions.
Speed, cost, and quality trade-offs
Streaming access and the WebSocket interface make Hy-ASR-3.0-Preview appropriate when an application needs transcription while a person is speaking. The catalog assessment gives it a relatively high editorial speed score, reflecting its real-time positioning, and a favorable editorial cost score. Those scores are comparative estimates rather than vendor measurements.
Pricing is usage-based through Tencent Cloud speech-recognition or related service billing. The reviewed documentation does not provide a verified standalone per-minute or per-request price for the model. Tencent associates the real-time model with its large-model 2.0 billing scheme, but the final charge depends on the selected Tencent Cloud product and workflow. Buyers should check the applicable service pricing before estimating costs; there is no reliable single price that can be applied to every Hy-ASR-3.0-Preview integration.
The central trade-off is specialization versus flexibility. A speech-recognition model can be faster and more economical for transcription than a general multimodal model, especially when the desired result is text with timestamps. In exchange, it lacks general reasoning, coding, tool use, speaker diarization, and unrestricted audio-session behavior. A broader speech or multimodal service may be more appropriate when those features matter more than the model’s Chinese dialect and mixed-language focus.
When to choose this model
Hy-ASR-3.0-Preview is a sensible choice when the main requirement is Tencent Cloud speech transcription involving Mandarin, English, Chinese dialects, or mixed Chinese-English audio. Its strongest practical use cases include:
- Real-time captions for Chinese-language meetings, broadcasts, and livestreams.
- Short voice commands where the application needs text quickly.
- Call, meeting, and interview transcription, provided audio is segmented or processed through an appropriate asynchronous workflow.
- Subtitle extraction from audio and video through Tencent Cloud media-processing services.
- Dialect-aware transcription for supported regional varieties.
- Enterprise audio containing product names or other specialized terms that can benefit from hotword injection.
It is less suitable when the application requires speaker-separated transcripts, built-in VAD, unrestricted long-session streaming, extensive vocabulary customization, general conversation, code generation, or autonomous tool use. In those cases, a different service type—or an additional post-processing pipeline—may be more appropriate.
Overall assessment
Hy-ASR-3.0-Preview is best understood as a focused Tencent Cloud ASR option rather than a general-purpose AI model. Its distinguishing value is the combination of Mandarin, English, mixed-language recognition, and support for 20 Chinese dialects, together with contextual correction and real-time access. The preview status, one-minute real-time session limit, 16 kHz mono PCM requirement, and missing diarization and VAD features are equally important when evaluating it.
Choose it when accurate, fast transcription for supported Chinese-language audio is the priority and Tencent Cloud is an acceptable deployment environment. Evaluate another option when the project needs a fully featured long-form transcription system, speaker attribution, extensive audio controls, or capabilities beyond converting speech into text.

