Hunyuan-Large

Hunyuan-Large

by Tencent AI · Available as an open-weight release; legacy for the historical Tencent Cloud hunyuan-large API identifier

Tencent Hunyuan-Large is a 389-billion-parameter open-weight mixture-of-experts language model with approximately 52 billion active parameters. It supports Chinese and English text tasks, reasoning, mathematics, coding, and long-context analysis, with pre-trained, instruction-tuned, and FP8 checkpoints. Its main trade-offs are substantial infrastructure requirements, no verified exact hosted pricing, and a legacy historical cloud API identifier.

Text Reasoning Coding
Tencent Hunyuan-Large is a large open-weight language model designed for organizations that want substantial model capacity, long-context processing, and the ability to run or fine-tune the model themselves. Its mixture-of-experts architecture contains 389 billion total parameters but activates approximately 52 billion for each token, helping separate overall capacity from per-token computation. Tencent released multiple checkpoints and supporting deployment resources, including vLLM and TensorRT-LLM guidance. The main trade-off is practical: Hunyuan-Large offers a very large model footprint and strong reported results across Chinese, English, mathematics, coding, and reasoning tasks, but it requires significant accelerator memory and distributed infrastructure. Its historical Tencent Cloud API identifier is now treated as legacy rather than a reliable indicator of a current hosted endpoint.
Outputs

What Hunyuan-Large can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
4/10 Speed
3/10 Cost efficiency
Specifications

Technical details

Model family Hunyuan-Large
Model type General Purpose
Context window 256K tokens
Release date 2024-11-05
Status Available as an open-weight release; legacy for the historical Tencent Cloud hunyuan-large API identifier
Knowledge cutoff notes

Tencent's official technical report and model documentation do not state a definitive knowledge cutoff for the exact Hunyuan-Large checkpoints.

Model notes

Hunyuan-Large is the official family name for Tencent's Hunyuan-MoE-A52B release. Tencent published Hunyuan-A52B-Pretrain, Hunyuan-A52B-Instruct, and Hunyuan-A52B-Instruct-FP8 checkpoints. The pre-trained checkpoint supports up to 256K tokens; Tencent states that the instruction-tuned model supports up to 128K tokens. The model has 389 billion total parameters and approximately 52 billion active parameters. Public source and checkpoint repositories remain available, but Tencent Cloud documentation maps the historical hunyuan-large API identifier to hy3-preview in its migration guidance. No current exact per-token price was verified for the open-weight checkpoints. Editorial scores are comparative estimates rather than vendor specifications.

Model guide

Tencent Hunyuan-Large: An Open-Weight 389B Mixture-of-Experts Model

Tencent Hunyuan-Large is an open-weight mixture-of-experts language model with 389 billion total parameters and approximately 52 billion active parameters per token. It targets Chinese and English text generation, reasoning, mathematics, coding, question answering, and long-context analysis, with a pre-trained context length of up to 256K tokens. Tencent released pre-trained, instruction-tuned, and FP8 variants, but the model is demanding to deploy and has no verified current per-token price for its open-weight checkpoints.

What is Hunyuan-Large?

Hunyuan-Large is Tencent's large-scale open-weight language model family, released on November 5, 2024. It is primarily a text-generation model for Chinese and English workloads, including instruction following, reasoning, mathematics, coding, question answering, and analysis of long documents.

The model uses a mixture-of-experts (MoE) architecture. Instead of evaluating every parameter for every token, an MoE model routes each token through selected expert networks. Hunyuan-Large has 389 billion total parameters and activates approximately 52 billion parameters per token. The total figure describes the model's available capacity, while the active-parameter figure is more relevant to the amount of computation used for an individual token. It should not be interpreted as making the model inexpensive to operate: the full checkpoint and its serving infrastructure remain very large.

Tencent released pre-trained, instruction-tuned, and FP8 variants under the Tencent Hunyuan Community License. The official repository includes model code, inference resources, deployment guidance, and fine-tuning material. Checkpoints are available through Tencent-hosted resources and Hugging Face.

Where Hunyuan-Large fits in Tencent's catalog

Hunyuan-Large is best understood as an open-weight research and self-hosting release rather than simply a current managed API model. It gives technical teams access to model checkpoints and the surrounding tools needed to experiment with local or distributed inference and fine-tuning.

Tencent Cloud documentation previously used the hunyuan-large identifier for an API-accessible model. Tencent's later migration documentation maps that historical identifier to hy3-preview. Consequently, the old API name should not be treated as the current canonical hosted endpoint for the exact open-weight Hunyuan-Large release. The open checkpoints and source repository remain the clearest way to identify and use this particular model.

This distinction matters when evaluating availability. A team can verify that the open-weight release exists and can obtain its checkpoints, but it should not assume that an old cloud API example represents a currently supported, separately priced Hunyuan-Large endpoint.

Architecture and context limits

The pre-trained Hunyuan-Large checkpoint supports sequences of up to 256K tokens, according to Tencent's technical documentation. A token is a unit of text used by a language model; the context limit is the amount of input and generated conversation history that the model can consider within one request. A 256K-token window is suitable for unusually large documents, repositories, or collections of related material, although the practical quality and memory requirements of very long requests still depend on the serving configuration.

The instruction-tuned checkpoint has a documented context limit of up to 128K tokens rather than 256K. Users therefore need to select the checkpoint based on the task: the pre-trained release has the longer stated context capacity, while the instruction-tuned release is intended to follow user instructions more directly.

Tencent describes grouped-query attention and cross-layer attention as part of the architecture used to reduce key-value cache memory requirements. These techniques can make long-context serving more manageable than an otherwise comparable architecture, but they do not remove the hardware demands of a model with this parameter scale. No maximum output-token limit for the exact open-weight release was verified in the supplied specifications.

Capabilities and reported evaluations

Hunyuan-Large is intended for general-purpose language work rather than a single narrow application. Its documented target areas include:

  • Chinese and English text generation and understanding
  • Instruction following and question answering
  • Mathematical problem solving
  • Reasoning and commonsense tasks
  • Source-code generation and coding assistance
  • Long-context analysis

Tencent's published evaluations cover English and Chinese benchmarks including MMLU, MMLU-Pro, CMMLU, C-Eval, BBH, GSM8K, MATH, HumanEval, and MBPP. The supplied research describes particularly strong reported performance on Chinese-language tasks, mathematics, commonsense understanding, and general language benchmarks. These are provider-reported evaluation results, not a guarantee that the model will outperform alternatives on every prompt, language, dataset, or deployment setup.

For coding, the model's published evaluation coverage includes HumanEval and MBPP. That supports using it for code generation and programming problem experiments, but it does not establish reliability for production code without testing, review, and application-specific validation. Likewise, reported mathematics and reasoning performance indicates intended capability rather than a guarantee of correct answers on difficult or unfamiliar problems.

Deployment and fine-tuning

Hunyuan-Large is aimed at teams with the infrastructure to operate a very large model. Tencent provides deployment paths based on vLLM and TensorRT-LLM, along with examples for local inference and distributed serving. The checkpoints are also compatible with Hugging Face-style workflows, and Tencent publishes DeepSpeed-based training resources for fine-tuning.

The FP8 variant can reduce memory requirements compared with a full-precision deployment. Quantization changes how model values are represented so that serving can use less memory, although it may involve trade-offs in hardware compatibility, throughput, or output quality. Even with FP8 or other optimization techniques, practical deployment remains substantially more demanding than serving a smaller dense model or a smaller MoE model.

The model data indicates streaming support in compatible serving setups. This should be understood as a serving capability rather than evidence of a current official hosted API with a particular streaming protocol. Teams should verify the behavior of the selected vLLM, TensorRT-LLM, or other deployment stack.

Modalities, tools, and output formats

Hunyuan-Large is a text-in, text-out language model. It does not provide native image, audio, video, music, embedding, speech, or action output in the supplied specifications. It is therefore appropriate for text generation and analysis, not for directly creating images or audio or for replacing a dedicated embedding model.

The open release does not establish a separate provider-native web-search tool. Function or tool-use support is also not verified as a distinct official capability for this exact open-weight model. Applications can potentially connect model output to external software, but that is an application-level integration and should not be presented as a guaranteed built-in tool interface.

A provider-native JSON-schema or structured-output mode is not established in the supplied research. The model may be prompted to produce JSON text, but that is different from a verified constrained decoding or schema-enforcement feature. Systems that require guaranteed machine-readable output should validate responses and implement their own error handling.

Pricing and cost trade-offs

No official per-token input or output price was verified for the exact open-weight Hunyuan-Large checkpoints. The open release should therefore not be described as having a confirmed standard API price. The financial cost is instead largely determined by hardware, hosting, electricity, storage, engineering, and operational requirements.

Hunyuan-Large is not a natural choice for low-cost or low-latency workloads. Its MoE routing reduces the active parameters per token compared with evaluating all 389 billion parameters, but the total checkpoint size and distributed serving requirements remain significant. FP8 and deployment optimizations can improve the infrastructure profile, yet smaller models will generally be easier to host and faster to start, scale, and operate.

The editorial scores associated with this model rate reasoning and coding relatively highly, while rating speed and cost less favorably. Those scores are comparative editorial estimates, not Tencent specifications or benchmark results. They summarize the practical trade-off suggested by the model's size and published capabilities.

Main strengths and limitations

Strengths

  • Large model capacity: 389 billion total parameters and approximately 52 billion active parameters provide a substantial MoE architecture for demanding language tasks.
  • Long context: the pre-trained checkpoint supports up to 256K tokens, while the instruction-tuned checkpoint supports up to 128K.
  • Chinese and English coverage: the model is designed for both languages, with Tencent reporting strong results on several Chinese-language evaluations.
  • Broad task coverage: the release targets reasoning, mathematics, coding, question answering, instruction following, and document analysis.
  • Self-hosting flexibility: public checkpoints, source code, fine-tuning resources, vLLM guidance, and TensorRT-LLM guidance support experimentation outside a single managed endpoint.

Limitations

  • High infrastructure requirements: the model is considerably more demanding than smaller language models, even when using FP8 or other optimizations.
  • Unclear hosted pricing: no verified current per-token price is available for the exact open-weight checkpoints.
  • Legacy API identity: the historical hunyuan-large Tencent Cloud identifier has been mapped to hy3-preview in migration documentation.
  • No native non-text generation: it is not an image, audio, video, embedding, or speech model.
  • Unverified structured and tool interfaces: native JSON-schema output, web search, and official function calling are not established by the supplied research.
  • Different checkpoint limits: the 256K context figure applies to the pre-trained checkpoint; the instruction-tuned checkpoint is documented at up to 128K.

When to choose Hunyuan-Large

Choose Hunyuan-Large when the priority is access to a very large open-weight model for Chinese and English text work, long-context research, mathematics, coding, or experimentation with self-hosted inference and fine-tuning. It is particularly relevant when an organization can operate distributed accelerator infrastructure and wants control over checkpoints, deployment, and adaptation rather than relying only on a small managed model endpoint.

Another option is likely more appropriate when the main requirements are low cost, fast responses, simple deployment, or a clearly documented current API with predictable per-token billing. Smaller dense or sparse models will generally be easier to run for routine chat, lightweight extraction, and high-volume applications. A dedicated multimodal model is preferable for image, audio, or video input and output, while a model with verified structured-output or tool-calling support is preferable when reliable machine-readable actions are central to the application.

Hunyuan-Large is consequently best viewed as a high-capacity open-weight text model with substantial engineering requirements. Its appeal lies in model scale, long context, language coverage, and deployment control; its costs lie in hardware, operational complexity, uncertain hosted pricing, and the lack of verified native capabilities outside standard text generation.


Answers to Frequently Asked Questions

Is Hunyuan-Large available through a current Tencent Cloud API with standard pricing?
No verified current per-token price was established for the exact open-weight Hunyuan-Large checkpoints. Tencent Cloud previously used the identifier `hunyuan-large`, but migration documentation maps that historical identifier to `hy3-preview`. The open checkpoints and official repository are the clearest way to identify and use this specific release.
Can Hunyuan-Large be self-hosted and fine-tuned?
Yes. Tencent provides model checkpoints, source code, inference resources, deployment guidance for vLLM and TensorRT-LLM, and DeepSpeed-based fine-tuning resources. The model is intended for teams with substantial distributed accelerator, storage, and operational infrastructure.
What is Tencent Hunyuan-Large?
Tencent Hunyuan-Large is an open-weight Chinese and English language model released on November 5, 2024. It uses a mixture-of-experts architecture with 389 billion total parameters and approximately 52 billion active parameters per token, supporting instruction following, reasoning, mathematics, coding, question answering, and long-document analysis.
How many tokens can Hunyuan-Large handle?
The pre-trained Hunyuan-Large checkpoint supports sequences of up to 256K tokens. The instruction-tuned checkpoint has a documented context limit of up to 128K tokens, so the applicable limit depends on which checkpoint is deployed.


Sources 6
Provider

About Tencent AI