Hy

Hy3

by Tencent AI · Current; open-weight and available through Tencent Cloud TokenHub

Tencent Hy3 is an open-weight 295B mixture-of-experts model with 21B active parameters, a 256K context window, and 128K maximum output. It targets coding, long-context reasoning, structured workflows, and tool-using agents, with TokenHub access and Apache 2.0 model weights.

Text Reasoning Coding
Tencent Hy3 is an open-weight reasoning and agent model built for coding, long-context analysis, structured workflows, and multi-step tool use. Its mixture-of-experts design contains 295 billion total parameters but activates 21 billion for a given inference, helping it target advanced reasoning while keeping the active computation lower than the total model size. Hy3 supports up to 256K tokens of context and is available through Tencent Cloud TokenHub as well as downloadable model weights.
Outputs

What Hy3 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Fine-tuning Structured output Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Hy
Model type Reasoning
Context window 256K tokens
Maximum output 128K tokens
Release date 2026-07-06
Status Current; open-weight and available through Tencent Cloud TokenHub
Knowledge cutoff notes

No authoritative first-party knowledge-cutoff date was identified for the exact Hy3 model.

Model notes

Hy3 uses a mixture-of-experts architecture with 295B total parameters, 21B active parameters, and 3.8B MTP-layer parameters. The model has 192 experts with top-8 activation and supports BF16 inference. Tencent's official repository states that the weights are released under the Apache 2.0 license and provides deployment and fine-tuning guidance for vLLM and SGLang. Tencent Cloud lists Hy3 with 256K context, 192K maximum input, and 128K maximum output. The canonical API model ID is hy3. The older hy3-preview identifier is routed to or migrated toward Hy3 in relevant Tencent Cloud services and should not be treated as a separate current model entity. Pricing is Tencent Cloud TokenHub pricing and may vary by region, service, contract, or billing plan. Reasoning, coding, speed, and cost values are editorial comparative estimates rather than provider-issued scores.

Cost

Model pricing

Input CNY 1 per 1 million input tokens; cached input CNY 0.25 per 1 million tokens on Tencent Cloud TokenHub
Output CNY 4 per 1 million output tokens on Tencent Cloud TokenHub
Model guide

Tencent Hy3: Open-Weight Reasoning for Coding and Agent Workflows

Tencent Hy3 is a 295-billion-parameter mixture-of-experts reasoning model with 21 billion active parameters, a 256K-token context window, 128K maximum output, tool use, structured outputs, coding support, and Apache 2.0 open-weight licensing.

What is Tencent Hy3?

Hy3 is a reasoning-focused large language model from Tencent's Hunyuan model family. It is designed to handle tasks that require more than a short answer, including code generation, long-document analysis, planning, structured responses, and interactions with external tools. In practical terms, Hy3 is aimed at applications where the model must break down a problem, keep track of substantial context, and complete several connected steps.

The model is provided in two main forms. Developers can call the current hy3 model through Tencent Cloud TokenHub, or download the model weights from Tencent's official Hy3 repository for supported self-managed deployments. Tencent states that the weights use the Apache 2.0 license. The older hy3-preview identifier should not be treated as a separate current model: relevant Tencent Cloud services route or migrate it toward the current Hy3 offering.

Hy3 is a current model in Tencent's catalog rather than a consumer chatbot product by itself. It is also distinct from Yuanbao, Tencent's consumer AI assistant. Hy3 may be integrated into Tencent products, but this page concerns the underlying model and its developer-facing deployment options.

Architecture and capacity

Hy3 uses a mixture-of-experts, or MoE, architecture. The model has 295 billion total parameters, while approximately 21 billion are active for an individual inference. Parameters are the learned numerical values used by a model to process input and generate output; in an MoE system, specialized subsets called experts can be selected instead of running every parameter for every token.

The official model information also describes 192 experts with top-8 activation, meaning that eight experts are selected for relevant processing at a routing step. Hy3 includes 3.8 billion parameters in its multi-token prediction layer and supports BF16 inference. These details matter most to teams planning local or private deployment: although the active parameter count is lower than the total, a 295B open-weight model still requires substantial computing, memory, and engineering resources.

SpecificationVerified detail
Model familyHy
ProviderTencent
Total parameters295 billion
Active parameters21 billion
ArchitectureMixture of experts
Context window256,000 tokens
Maximum input listed by Tencent Cloud192,000 tokens
Maximum output128,000 tokens
Weights licenseApache 2.0
Canonical API model IDhy3

The 256K context figure is the model's overall context capacity. Tencent Cloud separately lists a maximum input of 192K tokens and a maximum output of 128K tokens. These limits should not be added together for a single request; the service's effective input, output, and total-context rules apply when a request is submitted.

Reasoning, coding, and tool use

Hy3's main purpose is complex text-based work rather than media generation. Its reasoning orientation makes it suitable for tasks such as decomposing a software problem, comparing implementation choices, following a long set of instructions, or producing a multi-stage plan. The research describes reasoning modes and positions the model for agent workflows, where an application repeatedly asks the model to interpret information, decide on an action, use a tool, and continue from the result.

Coding is one of Hy3's strongest intended uses. It can be used for code generation, code explanation, debugging assistance, repository-level analysis, and programming agents that need to call external development tools. A coding workflow might provide a long project specification, ask Hy3 to identify the relevant files, call a search or test tool, and then generate a patch or explain the remaining failure. The model's 256K context is particularly relevant when a task involves extensive documentation, logs, or multiple source files.

Tool use and function calling are supported in the supplied model information. This allows an application to expose defined operations such as search, database lookup, file retrieval, or task execution. Hy3 can select and populate a tool call, while the surrounding application remains responsible for actually running the operation and validating its result. Tool support therefore does not mean that the hosted model can automatically access every system or website.

Structured output is also listed as supported. This is useful when a program needs a predictable object or schema rather than free-form prose, for example a list of extracted fields, a classification result, or a workflow action. A separate JSON-mode capability was not independently verified, so structured output should not automatically be interpreted as a distinct JSON mode.

Supported modalities

Hy3 is a text model in the supplied specifications. It accepts text input and produces text output. Native image, audio, and video input are not listed, and the model does not directly generate images, audio, video, music, or speech. Applications can still pass descriptions of media or use separate recognition and generation systems around Hy3, but that would be an application-level workflow rather than a native Hy3 modality.

This limitation helps define the model's position. Hy3 is appropriate for reasoning over text, code, instructions, tool results, and long documents. A user who needs image understanding, image generation, voice interaction, or video creation should select a model or product that explicitly supports those media types instead of assuming that Hy3 inherits all capabilities associated with Tencent's broader AI ecosystem.

Pricing and access

Tencent Cloud TokenHub lists Hy3 at CNY 1 per 1 million input tokens and CNY 4 per 1 million output tokens. Cached input is listed at CNY 0.25 per 1 million tokens. These are usage-based TokenHub prices, not a consumer subscription price, and Tencent notes that actual pricing can vary by region, service, contract, or billing arrangement.

Input and output tokens are billed differently, so workloads that generate long answers can cost more than workloads that mainly submit large prompts. Cached input can reduce the cost of repeated context when the service recognizes eligible cached content. Teams should also account for the infrastructure and operational cost of self-hosting the open weights; the Apache 2.0 license does not make large-scale inference hardware free.

Hy3 is listed as available through Tencent Cloud TokenHub, with the model ID hy3. Tencent's official repository provides deployment and fine-tuning guidance for vLLM and SGLang. The research confirms fine-tuning support in the model data, but it does not specify a single hardware configuration, throughput guarantee, or universally applicable fine-tuning recipe.

Main strengths and limitations

Where Hy3 is strong

  • Long-context work: The 256K-token context window and 192K maximum listed input support extensive documents, codebases, specifications, and tool histories.
  • Reasoning-oriented workflows: The model is designed for multi-step analysis, planning, and agent-style task execution rather than only short conversational responses.
  • Coding and automation: Coding support, tool use, function calling, and structured output fit software agents and business workflows that need machine-readable results.
  • Open-weight deployment: Apache 2.0 weights provide more control over deployment and integration than a hosted-only model, subject to the engineering requirements of running a model of this scale.
  • Token-based pricing: The listed input price is relatively low for workloads that need Hy3's reasoning and context capabilities, although the total cost depends on prompt and output volume.

Where Hy3 is limited

  • High deployment requirements: A 295B-parameter MoE model is not a lightweight local model. Organizations without suitable infrastructure may find TokenHub more practical than self-hosting.
  • Text-only native interface: Hy3 does not provide native image, audio, or video input or output in the supplied specifications.
  • Potential latency and cost trade-offs: Reasoning depth, long prompts, and large generated outputs can increase response time and token usage compared with smaller, speed-first models.
  • No verified knowledge-cutoff date: An authoritative cutoff date for the exact Hy3 model was not identified. Time-sensitive facts should therefore be checked with a current retrieval or web-search workflow rather than assumed to be known by the model.
  • Service-specific behavior: Limits, pricing, availability, and supported interfaces can differ between TokenHub and self-managed deployments. The cloud documentation should be checked before production integration.

The scores sometimes used to describe Hy3's reasoning, coding, speed, or cost are editorial comparative estimates in the supplied data, not provider-issued benchmark results. They should be treated as directional evaluations rather than measured guarantees.

When to choose Tencent Hy3

Choose Hy3 when the central requirement is a text-based reasoning or coding system that can work across a large context, call tools, and return structured results. It is a strong candidate for coding agents, repository analysis, document-heavy automation, planning systems, and workflows where several model-tool exchanges are needed. The open-weight option is also relevant to teams that need more control over deployment than a hosted-only API permits.

TokenHub is likely the simpler starting point when a team wants to test Hy3 without operating its own large-model infrastructure. Self-managed weights become more attractive when deployment control, customization, data handling, or integration with an existing inference stack outweighs the operational complexity.

Another type of model may be more appropriate when the task is primarily simple classification, short-form extraction, high-volume low-latency generation, or inexpensive chat. A smaller model can be faster and easier to deploy for those workloads. A multimodal model is the better choice for direct image, audio, or video understanding and generation. Finally, applications that depend on verified current information should pair Hy3 with a retrieval or web-search tool; the supplied research does not establish a knowledge-cutoff date for Hy3.

Bottom line

Tencent Hy3 is best understood as a large, open-weight, text-only reasoning model for demanding coding and agent workflows. Its defining practical features are the 256K context window, 128K maximum output, tool and structured-output support, Apache 2.0 weights, and availability through TokenHub. Those capabilities come with meaningful infrastructure and latency trade-offs, so Hy3 is most compelling when long-context reasoning, coding, and workflow control matter more than the smallest footprint or fastest possible response.


Answers to Frequently Asked Questions

What is Tencent Hy3 designed for?
Tencent Hy3 is a reasoning-focused, text-only large language model designed for coding, long-document analysis, planning, structured outputs, and agent workflows that use external tools.
What are Tencent Hy3's context window and model specifications?
Hy3 uses a mixture-of-experts architecture with 295 billion total parameters and approximately 21 billion active parameters. It supports a 256,000-token context window, with Tencent Cloud listing a maximum input of 192,000 tokens and a maximum output of 128,000 tokens.
Does Tencent Hy3 support image, audio, or video generation?
No. Hy3 is specified as a text-only model that accepts and produces text. Native image, audio, and video input or output are not listed, so those tasks require separate specialized models or application-level integrations.
How much does Tencent Hy3 cost through TokenHub?
Tencent Cloud TokenHub lists Hy3 at CNY 1 per 1 million input tokens, CNY 4 per 1 million output tokens, and CNY 0.25 per 1 million cached input tokens. Actual pricing may vary by region, service, contract, or billing arrangement.
How can developers access and deploy Hy3?
Developers can use the current hy3 model through Tencent Cloud TokenHub or download the Apache 2.0-licensed weights from Tencent's official Hy3 repository. The repository provides deployment and fine-tuning guidance for vLLM and SGLang.


Sources 7
Provider

About Tencent AI