What is Tencent Hy3?
Hy3 is a reasoning-focused large language model from Tencent's Hunyuan model family. It is designed to handle tasks that require more than a short answer, including code generation, long-document analysis, planning, structured responses, and interactions with external tools. In practical terms, Hy3 is aimed at applications where the model must break down a problem, keep track of substantial context, and complete several connected steps.
The model is provided in two main forms. Developers can call the current hy3 model through Tencent Cloud TokenHub, or download the model weights from Tencent's official Hy3 repository for supported self-managed deployments. Tencent states that the weights use the Apache 2.0 license. The older hy3-preview identifier should not be treated as a separate current model: relevant Tencent Cloud services route or migrate it toward the current Hy3 offering.
Hy3 is a current model in Tencent's catalog rather than a consumer chatbot product by itself. It is also distinct from Yuanbao, Tencent's consumer AI assistant. Hy3 may be integrated into Tencent products, but this page concerns the underlying model and its developer-facing deployment options.
Architecture and capacity
Hy3 uses a mixture-of-experts, or MoE, architecture. The model has 295 billion total parameters, while approximately 21 billion are active for an individual inference. Parameters are the learned numerical values used by a model to process input and generate output; in an MoE system, specialized subsets called experts can be selected instead of running every parameter for every token.
The official model information also describes 192 experts with top-8 activation, meaning that eight experts are selected for relevant processing at a routing step. Hy3 includes 3.8 billion parameters in its multi-token prediction layer and supports BF16 inference. These details matter most to teams planning local or private deployment: although the active parameter count is lower than the total, a 295B open-weight model still requires substantial computing, memory, and engineering resources.
| Specification | Verified detail |
|---|---|
| Model family | Hy |
| Provider | Tencent |
| Total parameters | 295 billion |
| Active parameters | 21 billion |
| Architecture | Mixture of experts |
| Context window | 256,000 tokens |
| Maximum input listed by Tencent Cloud | 192,000 tokens |
| Maximum output | 128,000 tokens |
| Weights license | Apache 2.0 |
| Canonical API model ID | hy3 |
The 256K context figure is the model's overall context capacity. Tencent Cloud separately lists a maximum input of 192K tokens and a maximum output of 128K tokens. These limits should not be added together for a single request; the service's effective input, output, and total-context rules apply when a request is submitted.
Reasoning, coding, and tool use
Hy3's main purpose is complex text-based work rather than media generation. Its reasoning orientation makes it suitable for tasks such as decomposing a software problem, comparing implementation choices, following a long set of instructions, or producing a multi-stage plan. The research describes reasoning modes and positions the model for agent workflows, where an application repeatedly asks the model to interpret information, decide on an action, use a tool, and continue from the result.
Coding is one of Hy3's strongest intended uses. It can be used for code generation, code explanation, debugging assistance, repository-level analysis, and programming agents that need to call external development tools. A coding workflow might provide a long project specification, ask Hy3 to identify the relevant files, call a search or test tool, and then generate a patch or explain the remaining failure. The model's 256K context is particularly relevant when a task involves extensive documentation, logs, or multiple source files.
Tool use and function calling are supported in the supplied model information. This allows an application to expose defined operations such as search, database lookup, file retrieval, or task execution. Hy3 can select and populate a tool call, while the surrounding application remains responsible for actually running the operation and validating its result. Tool support therefore does not mean that the hosted model can automatically access every system or website.
Structured output is also listed as supported. This is useful when a program needs a predictable object or schema rather than free-form prose, for example a list of extracted fields, a classification result, or a workflow action. A separate JSON-mode capability was not independently verified, so structured output should not automatically be interpreted as a distinct JSON mode.
Supported modalities
Hy3 is a text model in the supplied specifications. It accepts text input and produces text output. Native image, audio, and video input are not listed, and the model does not directly generate images, audio, video, music, or speech. Applications can still pass descriptions of media or use separate recognition and generation systems around Hy3, but that would be an application-level workflow rather than a native Hy3 modality.
This limitation helps define the model's position. Hy3 is appropriate for reasoning over text, code, instructions, tool results, and long documents. A user who needs image understanding, image generation, voice interaction, or video creation should select a model or product that explicitly supports those media types instead of assuming that Hy3 inherits all capabilities associated with Tencent's broader AI ecosystem.
Pricing and access
Tencent Cloud TokenHub lists Hy3 at CNY 1 per 1 million input tokens and CNY 4 per 1 million output tokens. Cached input is listed at CNY 0.25 per 1 million tokens. These are usage-based TokenHub prices, not a consumer subscription price, and Tencent notes that actual pricing can vary by region, service, contract, or billing arrangement.
Input and output tokens are billed differently, so workloads that generate long answers can cost more than workloads that mainly submit large prompts. Cached input can reduce the cost of repeated context when the service recognizes eligible cached content. Teams should also account for the infrastructure and operational cost of self-hosting the open weights; the Apache 2.0 license does not make large-scale inference hardware free.
Hy3 is listed as available through Tencent Cloud TokenHub, with the model ID hy3. Tencent's official repository provides deployment and fine-tuning guidance for vLLM and SGLang. The research confirms fine-tuning support in the model data, but it does not specify a single hardware configuration, throughput guarantee, or universally applicable fine-tuning recipe.
Main strengths and limitations
Where Hy3 is strong
- Long-context work: The 256K-token context window and 192K maximum listed input support extensive documents, codebases, specifications, and tool histories.
- Reasoning-oriented workflows: The model is designed for multi-step analysis, planning, and agent-style task execution rather than only short conversational responses.
- Coding and automation: Coding support, tool use, function calling, and structured output fit software agents and business workflows that need machine-readable results.
- Open-weight deployment: Apache 2.0 weights provide more control over deployment and integration than a hosted-only model, subject to the engineering requirements of running a model of this scale.
- Token-based pricing: The listed input price is relatively low for workloads that need Hy3's reasoning and context capabilities, although the total cost depends on prompt and output volume.
Where Hy3 is limited
- High deployment requirements: A 295B-parameter MoE model is not a lightweight local model. Organizations without suitable infrastructure may find TokenHub more practical than self-hosting.
- Text-only native interface: Hy3 does not provide native image, audio, or video input or output in the supplied specifications.
- Potential latency and cost trade-offs: Reasoning depth, long prompts, and large generated outputs can increase response time and token usage compared with smaller, speed-first models.
- No verified knowledge-cutoff date: An authoritative cutoff date for the exact Hy3 model was not identified. Time-sensitive facts should therefore be checked with a current retrieval or web-search workflow rather than assumed to be known by the model.
- Service-specific behavior: Limits, pricing, availability, and supported interfaces can differ between TokenHub and self-managed deployments. The cloud documentation should be checked before production integration.
The scores sometimes used to describe Hy3's reasoning, coding, speed, or cost are editorial comparative estimates in the supplied data, not provider-issued benchmark results. They should be treated as directional evaluations rather than measured guarantees.
When to choose Tencent Hy3
Choose Hy3 when the central requirement is a text-based reasoning or coding system that can work across a large context, call tools, and return structured results. It is a strong candidate for coding agents, repository analysis, document-heavy automation, planning systems, and workflows where several model-tool exchanges are needed. The open-weight option is also relevant to teams that need more control over deployment than a hosted-only API permits.
TokenHub is likely the simpler starting point when a team wants to test Hy3 without operating its own large-model infrastructure. Self-managed weights become more attractive when deployment control, customization, data handling, or integration with an existing inference stack outweighs the operational complexity.
Another type of model may be more appropriate when the task is primarily simple classification, short-form extraction, high-volume low-latency generation, or inexpensive chat. A smaller model can be faster and easier to deploy for those workloads. A multimodal model is the better choice for direct image, audio, or video understanding and generation. Finally, applications that depend on verified current information should pair Hy3 with a retrieval or web-search tool; the supplied research does not establish a knowledge-cutoff date for Hy3.
Bottom line
Tencent Hy3 is best understood as a large, open-weight, text-only reasoning model for demanding coding and agent workflows. Its defining practical features are the 256K context window, 128K maximum output, tool and structured-output support, Apache 2.0 weights, and availability through TokenHub. Those capabilities come with meaningful infrastructure and latency trade-offs, so Hy3 is most compelling when long-context reasoning, coding, and workflow control matter more than the smallest footprint or fastest possible response.
Answers to Frequently Asked Questions
hy3 model through Tencent Cloud TokenHub or download the Apache 2.0-licensed weights from Tencent's official Hy3 repository. The repository provides deployment and fine-tuning guidance for vLLM and SGLang.
