Hy-ASR

Hy-ASR-3.0-Preview

by Tencent AI · Preview; internal-test availability

Tencent Hy-ASR-3.0-Preview is a preview speech-recognition model for Mandarin, English, mixed Chinese-English audio, and 20 Chinese dialects. It supports real-time WebSocket transcription, asynchronous media workflows, contextual correction, hotwords, timestamps, and subtitle extraction. Its main limitations are the one-minute real-time session limit, 16 kHz mono PCM requirement, unavailable diarization and VAD features, and unverified standalone pricing.

Text Reasoning Coding
Tencent Hy-ASR-3.0-Preview is a preview-generation speech-recognition model for converting spoken audio into text. Its main distinction is a focus on contextual accuracy across Mandarin, English, mixed-language speech, and a broad set of Chinese dialects. It can be accessed through Tencent Cloud real-time WebSocket recognition, media-processing services, and TokenHub synchronous transcription interfaces, although the preview service has important input and feature restrictions.
Outputs

What Hy-ASR-3.0-Preview can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

4/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Hy-ASR
Model type Other
Release date 2026-08-04
Status Preview; internal-test availability
Knowledge cutoff notes

Tencent's reviewed documentation does not publish a knowledge cutoff for this speech-recognition model. Its output is generated from supplied audio rather than from a documented static language-model knowledge date.

Model notes

Hy-ASR-3.0-Preview is a speech-recognition model rather than a general-purpose conversational model. Tencent documents support for Mandarin, English, and 20 Chinese dialects, contextual correction, hotword injection, and difficult acoustic conditions. The real-time preview interface is limited to audio of up to 60 seconds or one minute per session and 16 kHz mono PCM input. The internal-test version does not support speaker diarization, VAD, vocabulary replacement, or noise-threshold parameters. Real-time access uses the model identifier Hy-ASR-3.0-preview; TokenHub uses hy-asr-3.0-preview. Output can include full transcription, sentence timestamps, word-level timestamps in some interfaces, and subtitle files through applicable Tencent Cloud services. Editorial scores are comparative estimates for a speech-recognition model, not vendor benchmarks.

Cost

Model pricing

Input Usage-based Tencent Cloud ASR pricing; exact model price not verified in the reviewed official documentation
Output No separate output price; transcription is billed through the applicable Tencent Cloud ASR or large-model 2.0 pricing scheme
Model guide

Hy-ASR-3.0-Preview: Tencent’s Context-Aware ASR for Chinese Dialects and Mixed Speech

Tencent Hy-ASR-3.0-Preview is a preview automatic speech-recognition model for Mandarin, English, and 20 Chinese dialects. It combines speech recognition with Hy3-based language understanding to improve contextual correction, mixed Chinese-English transcription, hotword handling, and robustness in difficult audio. The model is designed for real-time and asynchronous transcription, subtitles, voice commands, and media-processing workflows rather than general text generation.

What Hy-ASR-3.0-Preview is

Hy-ASR-3.0-Preview is Tencent’s preview automatic speech-recognition (ASR) model. ASR systems analyze spoken audio and return a written transcription; this model is intended specifically for that task rather than for open-ended conversation, coding, image generation, or general text completion.

Within Tencent’s current product lineup, it sits in the Hunyuan speech-recognition offering and is exposed through several Tencent Cloud services. The preview and internal-test status means that availability and feature support are more limited than they would be for a mature, unrestricted production model. Tencent’s documentation identifies the real-time model as Hy-ASR-3.0-preview; TokenHub uses the lowercase identifier hy-asr-3.0-preview.

Languages and recognition strengths

The documented language coverage includes Mandarin Chinese, English, and 20 Chinese dialects or regional varieties. These include Cantonese, Northeastern Mandarin, Henan, Shaanxi, Chengdu, Chongqing, Wuhan, Guiyang, Qingdao, Jinan, Changsha, Hefei, Hebei, Kunming, Lanzhou, Yinchuan, Nanchang, Beijing, Sichuan, and Tianjin varieties.

This coverage makes the model more specialized than a speech-to-text service focused only on standard Mandarin. It is also intended to handle Chinese-English mixed speech, which is common in technical discussions, customer calls, livestreams, and workplace conversations. Tencent describes improvements in contextual disambiguation and homophone correction, meaning the model attempts to use surrounding words to choose a more plausible transcription instead of treating every sound in isolation.

Hotword injection is another documented capability. A hotword is a user-supplied term that the recognizer should pay particular attention to, such as a product name, company name, person’s name, medical term, or specialist vocabulary. This can be useful when ordinary language models might replace an uncommon term with a more familiar-sounding word.

Supported inputs and outputs

Hy-ASR-3.0-Preview accepts audio and produces text. It does not provide direct image, video, audio, music, embedding, or general structured-object output. Video can still be processed indirectly through Tencent Cloud media services when the service extracts or transcribes the audio track for subtitles.

The real-time preview interface is documented for 16 kHz, mono PCM audio. The preview real-time service supports audio of up to 60 seconds, or one minute, per recognition session. This is a significant operational constraint: an application handling a long meeting, call, or livestream must divide the recording into suitable segments or use an asynchronous or media-processing workflow where supported.

Output can include a full transcription and sentence timestamps. Some interfaces also provide word-level timestamps, while applicable Tencent Cloud services can produce subtitle files. The supplied documentation does not publish a language-model context window, maximum output-token value, or a single universal output limit across all interfaces. Limits therefore depend on the selected Tencent Cloud service and workflow.

Access methods and practical workflows

For live applications, the model can be selected through Tencent Cloud’s real-time speech-recognition WebSocket API. WebSocket is a persistent connection that allows audio to be sent incrementally and transcription results to be returned as recognition proceeds. This makes the model suitable for live captions, voice-command interfaces, and interactive audio applications.

For files and media operations, Tencent Cloud Media Processing supports Hy ASR 3.0 preview integration for use cases such as subtitle extraction and synchronous audio recognition. TokenHub provides a synchronous transcription endpoint using the lowercase model identifier. These interfaces are useful when the application does not need to display partial results immediately or when audio is already stored as part of a media-processing job.

The exact request format, authentication requirements, regional availability, and service-level limits belong to the individual Tencent Cloud interface rather than to the model alone. Developers should therefore select the service first—real-time recognition, media processing, or TokenHub—and then follow that service’s current documentation.

Preview limitations and unsupported features

The most important limitation is the one-minute real-time session restriction documented for the preview interface. It makes the model a better fit for short live segments or applications that can manage segmentation than for a single unrestricted audio stream.

The internal-test version does not support speaker diarization, voice activity detection (VAD), vocabulary replacement, or noise-threshold parameters. Speaker diarization identifies which person is speaking, so applications that require labels such as “Speaker 1” and “Speaker 2” should not assume that this preview model provides them. VAD detects when speech starts and stops; without the documented VAD feature, the surrounding application or service may need to manage silence handling separately.

Vocabulary replacement and noise-threshold controls are also unavailable in the internal-test version. This limits customization for specialized terminology and for applications that need precise control over how background noise affects recognition. Tencent does describe improved stability in noisy, whispered, and otherwise difficult acoustic conditions, but that provider claim should not be treated as a substitute for a published benchmark or a guarantee for every recording environment.

Reasoning, coding, and tool support

Hy-ASR-3.0-Preview is not a reasoning or coding model in the usual language-model sense. Its contextual understanding is used to improve transcription—for example, resolving homophones or preserving mixed Chinese-English phrases—not to solve multi-step problems or generate software.

The supplied model assessment assigns a low comparative coding score and does not identify tool or function-calling support. These are editorial evaluations and catalog classifications, not Tencent-published benchmarks. The model should be treated as a transcription component that can be called by an application, not as an agent that independently invokes tools, browses the web, writes code, or performs business actions.

Speed, cost, and quality trade-offs

Streaming access and the WebSocket interface make Hy-ASR-3.0-Preview appropriate when an application needs transcription while a person is speaking. The catalog assessment gives it a relatively high editorial speed score, reflecting its real-time positioning, and a favorable editorial cost score. Those scores are comparative estimates rather than vendor measurements.

Pricing is usage-based through Tencent Cloud speech-recognition or related service billing. The reviewed documentation does not provide a verified standalone per-minute or per-request price for the model. Tencent associates the real-time model with its large-model 2.0 billing scheme, but the final charge depends on the selected Tencent Cloud product and workflow. Buyers should check the applicable service pricing before estimating costs; there is no reliable single price that can be applied to every Hy-ASR-3.0-Preview integration.

The central trade-off is specialization versus flexibility. A speech-recognition model can be faster and more economical for transcription than a general multimodal model, especially when the desired result is text with timestamps. In exchange, it lacks general reasoning, coding, tool use, speaker diarization, and unrestricted audio-session behavior. A broader speech or multimodal service may be more appropriate when those features matter more than the model’s Chinese dialect and mixed-language focus.

When to choose this model

Hy-ASR-3.0-Preview is a sensible choice when the main requirement is Tencent Cloud speech transcription involving Mandarin, English, Chinese dialects, or mixed Chinese-English audio. Its strongest practical use cases include:

  • Real-time captions for Chinese-language meetings, broadcasts, and livestreams.
  • Short voice commands where the application needs text quickly.
  • Call, meeting, and interview transcription, provided audio is segmented or processed through an appropriate asynchronous workflow.
  • Subtitle extraction from audio and video through Tencent Cloud media-processing services.
  • Dialect-aware transcription for supported regional varieties.
  • Enterprise audio containing product names or other specialized terms that can benefit from hotword injection.

It is less suitable when the application requires speaker-separated transcripts, built-in VAD, unrestricted long-session streaming, extensive vocabulary customization, general conversation, code generation, or autonomous tool use. In those cases, a different service type—or an additional post-processing pipeline—may be more appropriate.

Overall assessment

Hy-ASR-3.0-Preview is best understood as a focused Tencent Cloud ASR option rather than a general-purpose AI model. Its distinguishing value is the combination of Mandarin, English, mixed-language recognition, and support for 20 Chinese dialects, together with contextual correction and real-time access. The preview status, one-minute real-time session limit, 16 kHz mono PCM requirement, and missing diarization and VAD features are equally important when evaluating it.

Choose it when accurate, fast transcription for supported Chinese-language audio is the priority and Tencent Cloud is an acceptable deployment environment. Evaluate another option when the project needs a fully featured long-form transcription system, speaker attribution, extensive audio controls, or capabilities beyond converting speech into text.


Answers to Frequently Asked Questions

Is Hy-ASR-3.0-Preview suitable for long recordings or general AI tasks?
It can be used for long recordings through asynchronous or media-processing workflows that segment or process the audio appropriately, but its real-time interface is limited to one-minute sessions. It is a focused transcription model, not a general reasoning, coding, multimodal, or tool-using AI system.
How can developers access Hy-ASR-3.0-Preview?
Developers can access the model through Tencent Cloud’s real-time speech-recognition WebSocket API, Tencent Cloud Media Processing for stored audio and video workflows, or TokenHub’s synchronous transcription endpoint. The exact authentication, request format, regional availability, and service limits depend on the selected interface.
What are the main limitations of Hy-ASR-3.0-Preview?
The real-time preview interface supports 16 kHz mono PCM audio and limits each recognition session to 60 seconds. The internal-test version does not support speaker diarization, voice activity detection (VAD), vocabulary replacement, or noise-threshold parameters.
What is Hy-ASR-3.0-Preview used for?
Hy-ASR-3.0-Preview is Tencent’s automatic speech-recognition model for converting spoken audio into text. It is designed for real-time captions, voice commands, meeting and call transcription, subtitle extraction, and other speech-to-text workflows.
Which languages and Chinese dialects does Hy-ASR-3.0-Preview support?
The model supports Mandarin Chinese, English, Chinese-English mixed speech, and 20 Chinese dialects or regional varieties, including Cantonese, Northeastern Mandarin, Henan, Shaanxi, Chengdu, Chongqing, Wuhan, Beijing, Sichuan, and Tianjin varieties.


Sources 6
Provider

About Tencent AI