AuK

AuK

by Tencent AI · Current open-source release

Tencent AuK is an open-weight 1.5B audio foundation model that uses natural-language instructions and optional audio context for speech generation, voice-conditioned TTS, speech editing, enhancement, and separation.

Speech Reasoning Coding
Tencent AuK is an open-source audio model designed to handle speech generation and editing through a unified instruction-based interface. Instead of functioning as a conventional text chatbot, AuK takes written instructions and, when needed, reference or source audio, then produces generated or transformed waveform audio. Its main appeal is the breadth of speech and audio operations available from one local model, while its main limitation is that it is not documented as a hosted general-purpose language or API model.
Outputs

What AuK can produce

Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family AuK
Model type Other
Release date 2026-09-09
Status Current open-source release
Knowledge cutoff notes

No explicit provider-published knowledge cutoff was identified. AuK is an audio generation and editing model rather than a conventional text language model, and its documented operation relies on instructions, optional audio context, and local model weights.

Model notes

AuK is an approximately 1.5B open-weight audio foundation model released by Tencent Hunyuan and collaborators. It accepts natural-language instructions and optional audio context and returns generated or transformed waveform audio. Supported task families include zero-shot and instructed TTS, speech-content editing, lyric editing, pitch, speed and volume editing, emotion and timbre transformation, de-accenting, nonverbal editing, whisper conversion, speech enhancement, speech separation, music separation, and target-speaker extraction. The official release also provides AuK-Flash as a separate distilled four-step checkpoint; AuK-Flash is not the same canonical model record. No first-party hosted API model ID, standard token pricing, or production cloud endpoint is documented in the reviewed release material. The repository documents local inference and fine-tuning. AuK is released under the MIT License.

Model guide

Tencent AuK: Open-Source Speech Generation and Audio Editing in One Model

Tencent AuK is an approximately 1.5B-parameter open-weight audio foundation model for instruction-based speech generation and editing. It combines natural-language instructions with optional audio context to support zero-shot and instructed text-to-speech, speech editing, enhancement, separation, and related audio transformation tasks.

What is Tencent AuK?

Tencent AuK is an approximately 1.5B-parameter open-weight foundation model from Tencent Hunyuan and its collaborators. It is built for speech generation, speech editing, and broader audio transformation rather than general-purpose text conversation. The model accepts a natural-language instruction and can optionally use audio context, such as a reference voice or source recording. Its output is generated or modified waveform audio.

This instruction-based design lets users describe the requested operation in words instead of selecting a separate narrowly defined model for every task. For example, an instruction can request speech in a reference voice, ask for a change in emotion or speaking speed, or identify an editing operation such as removing an accent. The exact quality of an output depends on the task, prompt, source audio, and local deployment setup; the supplied release material does not establish a universal quality guarantee for every supported operation.

Primary purpose and capabilities

AuK is intended to unify several speech and audio workflows. Its documented capability set includes both generation and transformation:

  • Zero-shot text-to-speech: generate speech using a reference recording without a separately trained speaker-specific model.
  • Instruction-based text-to-speech: generate speech from a written description of the desired voice or delivery.
  • Speech and lyric editing: modify spoken content or lyrics while preserving relevant characteristics of the source audio.
  • Performance adjustment: change pitch, speed, or volume.
  • Voice and delivery transformation: modify emotion or timbre, remove an accent, or convert speech into a whisper-like form.
  • Nonverbal-sound editing: work with sounds that are not ordinary spoken words.
  • Audio restoration and separation: enhance speech, separate speech sources, separate music and vocals, or extract a target speaker.

These functions make AuK more relevant to audio production, localization, accessibility, voice prototyping, and research than to ordinary text-generation workloads. The model can be considered a general audio foundation model within its intended domain, but it should not be interpreted as a general multimodal assistant simply because it accepts both text instructions and audio.

How the model is built

The reported architecture combines three main elements. A multimodal language model provides semantic conditioning, meaning it helps interpret the instruction and the intended content or transformation. An audio variational autoencoder, or audio VAE, represents speech, general audio, and music in a form the generative system can process. A hybrid rectified-flow Transformer then combines semantic and acoustic information before waveform reconstruction.

The Transformer uses dual-stream MMDiT blocks followed by single-stream DiT blocks. In practical terms, this architecture first keeps different information streams distinct while they are being processed and later combines them for generation. Users do not need to understand these components to run AuK, but they help explain why the model is designed for both language-guided instructions and detailed audio transformations rather than only conventional text-to-speech.

Inputs, outputs, and modalities

AuK supports text instructions and audio input. Audio is optional for tasks that can be completed from a written description, but it is important for reference-voice generation and source-audio editing. The documented input modalities are therefore text and audio; image and video input are not identified as supported features.

The model's direct output is audio, including speech and other waveform results. It does not provide text, images, video, embeddings, or structured records as its primary output. This distinction matters when evaluating AuK against language models: a system may use text to control AuK, but the model itself is intended to return audio rather than a textual answer.

AuK and the AuK-Flash variant

The official release describes two related checkpoints. AuK is the full base model intended for high-quality generation, while AuK-Flash is a separate distilled checkpoint designed for four-step inference. Distillation generally means that a smaller or more streamlined model is trained to reproduce the behavior of a larger or fuller system with fewer inference steps.

AuK-Flash is useful context when considering speed, but it should not be silently treated as the same model record. The canonical subject here is AuK. The supplied research does not provide a complete quality, memory, or latency comparison between the two checkpoints, so it would be inappropriate to claim that one is universally better. The documented distinction is that AuK is the full model and AuK-Flash is the separate four-step variant.

Main strengths and trade-offs

AuK's strongest differentiator is task coverage within speech and audio. A single open-weight model can address voice-conditioned text-to-speech, instruction-guided synthesis, editing, enhancement, and separation. That can simplify experimentation where several separate audio utilities would otherwise be required.

Its open-weight availability is another practical advantage. The model is released with source code and weights under the MIT License, and the official materials document local inference, Gradio, ComfyUI, Python usage, and fine-tuning workflows. This gives technical users more control over deployment and experimentation than a closed, hosted-only service would provide.

The trade-off is operational rather than only technical. Local deployment requires users to manage the runtime, hardware, model files, and audio-processing workflow themselves. The release material does not document a first-party metered hosted API, a standard cloud model identifier, token pricing, or service-level guarantees. Organizations that need a managed endpoint, predictable billing, or provider-operated availability may therefore prefer a hosted audio service, even if it offers a narrower feature set.

Limits and undocumented specifications

Several conventional language-model specifications do not apply cleanly to AuK. No context length or maximum output-token limit is documented in the supplied research. These omissions are expected for a waveform-generating audio model: audio duration, sampling and processing configuration, memory use, and inference settings are not naturally represented by a text token limit alone. The available material does not provide a universal maximum input duration or output duration, so those values should not be assumed.

AuK is not documented as having JSON mode, structured output, tool or function calling, web search, batch API access, or streaming output. It also is not presented as a reasoning or coding model. An application can place AuK inside a larger pipeline that uses other software for orchestration, prompting, or post-processing, but those surrounding capabilities should not be attributed to AuK itself.

The technical report notes that unconstrained editing requests can benefit from explicit task routing and prompt enhancement. This is a meaningful limitation for practical use: a vague instruction may not identify the intended operation clearly enough. Structured task descriptions, suitable source audio, and explicit editing goals are likely to produce a more reliable workflow than treating every request as an unrestricted natural-language command. This is an implementation consideration, not a published guarantee of output quality.

Pricing and availability

No input or output price is available for AuK because the supplied release information describes an open-weight local model rather than a metered hosted API. There is no documented subscription, per-request charge, token price, or official cloud endpoint to use as a pricing reference. Running costs instead depend on the user's hardware, storage, hosting arrangement, and engineering resources.

The model and supporting materials are available through Tencent Hunyuan's public GitHub repository, Hugging Face, and ModelScope. The official repository documents multiple ways to work with the model, including local Python inference, a Gradio interface, ComfyUI integration, and fine-tuning workflows. AuK is released under the MIT License according to the supplied research.

When to choose AuK

AuK is a good candidate when the priority is local control over speech or audio generation and editing, especially when several related operations are needed in the same project. Suitable uses include:

  • Prototyping a voice-controlled or reference-voice text-to-speech workflow.
  • Generating speech from descriptions of voice characteristics or delivery style.
  • Experimenting with pitch, speed, volume, emotion, timbre, accent, or whisper transformations.
  • Editing speech or lyrics with natural-language instructions.
  • Enhancing recordings or separating speech, vocals, music, or a target speaker.
  • Building a research or production pipeline that requires open weights, local execution, or fine-tuning.

Another option may be more appropriate when the project requires a managed API, published usage pricing, guaranteed service levels, text responses, coding assistance, web search, structured JSON, or image and video generation. A specialized hosted text-to-speech service may also be preferable when operational simplicity and predictable latency matter more than local customization. Conversely, a narrowly focused audio tool may be easier to control for one specific separation or enhancement task, while AuK is more attractive when the workflow spans multiple generation and editing functions.

Overall assessment

Tencent AuK is best understood as an open audio foundation model with an instruction interface, not as a general-purpose conversational model. Its approximately 1.5B-parameter design, optional audio context, and broad set of speech and audio operations make it particularly relevant to users who want one locally deployable system for synthesis and transformation. The full AuK checkpoint prioritizes the primary release's generation capability, while AuK-Flash provides a separate four-step inference option for users investigating faster operation.

The decision to use AuK should ultimately depend on deployment requirements. Its strengths are open access, local control, fine-tuning support, and coverage across speech generation and editing. Its limitations are the lack of documented hosted pricing and service guarantees, the absence of conventional language-model features, and the need to manage local inference and task-specific prompting. Within those boundaries, AuK offers a practical foundation for speech and audio experimentation without requiring the user to treat it as a chatbot or a conventional API language model.


Answers to Frequently Asked Questions

When should developers choose Tencent AuK?
Developers should consider AuK when they need local control, open weights, fine-tuning, and one system for multiple speech and audio tasks such as text-to-speech, voice transformation, editing, enhancement, or separation. A managed audio API or specialized tool may be preferable when predictable pricing, service guarantees, simpler deployment, or a single narrowly defined operation is more important.
Is Tencent AuK available through a hosted API, and how much does it cost?
The supplied release information describes AuK as an open-weight local model rather than a metered hosted API. No official cloud endpoint, subscription, token price, or per-request charge is documented. Running costs depend on the user's hardware, storage, hosting, and engineering resources.
What is the difference between AuK and AuK-Flash?
AuK is the full checkpoint intended for high-quality generation, while AuK-Flash is a separate distilled checkpoint designed for four-step inference. The available research does not establish a complete quality, memory, or latency comparison between the two variants.
What is Tencent AuK designed to do?
Tencent AuK is an approximately 1.5B-parameter open-weight audio foundation model designed for speech generation, speech and lyric editing, voice transformation, audio restoration, and source separation. It uses natural-language instructions and can optionally use reference or source audio.
What inputs and outputs does Tencent AuK support?
AuK accepts text instructions and optional audio input, such as a reference voice or source recording. Its primary output is generated or modified waveform audio, including speech and other audio results. It is not documented as an image, video, text, embedding, or structured-record generation model.


Sources 5
Provider

About Tencent AI