Hy4

Hy4 preview

by Tencent AI · Preview; currently available as an open-weight model and through Tencent products, Tencent Cloud TokenHub, and OpenRouter

Tencent Hy4 preview is an open-weight, text-only Mixture-of-Experts model with 770 billion total parameters, 49 billion active parameters per token, and a 1-million-token context window. It is designed for long-context coding agents, tool-use workflows, productivity automation, document analysis, game development, and scientific reasoning. The model supports streaming, structured outputs, fine-tuning, cache support, and vLLM or SGLang deployment, with TokenHub pricing of CNY 6 per million input tokens and CNY 18 per million output tokens. Its main trade-offs are preview status, potentially lengthy reasoning, demanding self-hosting requirements, and no image, audio, or video capabilities.

Text Reasoning Coding
Hy4 preview is Tencent’s current open-weight flagship language model for demanding text-based workloads. It combines a 1-million-token context window with tool calling, structured outputs, fine-tuning support, and deployment guidance for vLLM and SGLang. The model is aimed at coding agents, long documents, productivity automation, game development, and scientific reasoning rather than image, audio, or video tasks. Its large total parameter count and preview status make it a specialist option: potentially attractive when context capacity and complex reasoning matter more than maximum speed, simple deployment, or the predictability of a mature production model.
Outputs

What Hy4 preview can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Fine-tuning Structured output Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
6/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Hy4
Model type General Purpose
Context window 1M tokens
Maximum output 64K tokens
Release date 2026-08-28
Status Preview; currently available as an open-weight model and through Tencent products, Tencent Cloud TokenHub, and OpenRouter
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was found in the reviewed Tencent model announcement, model repository, or TokenHub documentation.

Model notes

Hy4 preview is Tencent's Mixture-of-Experts flagship model with 770B total parameters and 49B active parameters per token. The architecture has 78 backbone layers, 256 routed experts plus one shared expert in each MoE layer, and an additional MTP layer for speculative decoding. Tencent documents a 1M-token context window, 960K maximum input, and 64K maximum output through TokenHub. The model defaults to high reasoning mode and supports a no-think setting through the deployment chat template. Tencent provides official vLLM and SGLang deployment guidance, an OpenAI-compatible local API, a fine-tuning workflow, tool calling, structured outputs, and cache support. The model is released under Apache License 2.0. Tencent describes it as an early preview with known limitations including lengthy reasoning and excessive self-verification on complex tasks. Pricing is for Tencent Cloud TokenHub and may differ across regions, plans, or third-party gateways.

Cost

Model pricing

Input CNY 6 per 1M tokens; cached input CNY 0.3 per 1M tokens
Output CNY 18 per 1M tokens
Model guide

Hy4 preview: Tencent’s open-weight model for long-context reasoning and coding agents

Hy4 preview is Tencent’s 770-billion-parameter Mixture-of-Experts language model, designed for long-context reasoning, coding agents, tool use, productivity automation, and complex document or scientific tasks. It offers a 1-million-token context window, up to 64,000 output tokens through TokenHub, open-weight deployment options, and relatively low active parameters per token, but remains an early preview with text-only input and output.

What is Hy4 preview?

Hy4 preview is a text-focused large language model from Tencent. Tencent describes it as a flagship Mixture-of-Experts, or MoE, model. In an MoE system, the model contains many specialist neural-network components, but only a selected subset is activated for each token. This allows a model to have a very large overall capacity without using every parameter for every piece of text.

Hy4 preview has 770 billion total parameters and 49 billion active parameters per token. Those figures describe the model’s architecture, not a guarantee of benchmark performance. In practical terms, the design is intended to provide substantial reasoning and language capacity while reducing the amount of computation used for each generated token compared with a dense model of the same total size.

The model was released as a preview and is available as an open-weight model, through Tencent products and Tencent Cloud TokenHub, and through third-party gateways including OpenRouter according to the supplied research. Tencent also provides documentation for local or self-managed deployment. Because it is a preview, users should treat its behavior, interfaces, availability, and operational reliability as less settled than those of a mature production release.

Where Hy4 preview fits in Tencent’s lineup

Hy4 preview belongs to Tencent’s Hunyuan model family and represents a high-end general-purpose language model in that family. It is distinct from Yuanbao, Tencent’s consumer AI assistant, even though Tencent’s model generations may be integrated into Yuanbao and other Tencent products. Hy4 preview should therefore be evaluated as a model that can be deployed or accessed through developer-oriented channels, not as a consumer chat subscription.

Its position is also different from Tencent’s image, video, 3D, and other multimodal services. Hy4 preview is documented as a text-in, text-out model. It does not accept images, audio, or video as model inputs, and it does not directly generate those media types. That makes it a focused language and agent model rather than a single model for every modality in Tencent’s broader AI ecosystem.

Key specifications and limits

SpecificationHy4 preview
ProviderTencent
Model familyHy4
StatusPreview
ArchitectureMixture-of-Experts
Total parameters770 billion
Active parameters per token49 billion
Context window1,000,000 tokens
Maximum input through TokenHub960,000 tokens
Maximum output through TokenHub64,000 tokens
Input and output modalitiesText input and text output
LicenseApache License 2.0
Tool useSupported
StreamingSupported
Fine-tuningSupported
Structured outputsSupported

The advertised 1-million-token context window is the model’s headline capacity. Tencent’s TokenHub documentation specifies up to 960,000 input tokens and 64,000 output tokens in that overall context arrangement. A token is a small unit of text used by language models; the exact number of words represented by a token count varies by language and content. The practical implication is that Hy4 preview can work with unusually large collections of text, although context capacity alone does not guarantee that every detail will be handled perfectly.

Reasoning and coding capabilities

Hy4 preview is intended for tasks that require more than short-form text completion. The supplied research rates its reasoning and coding capabilities at 8 out of 10 as editorial evaluations, not as scores published by Tencent. Those ratings indicate an assessment that reasoning and programming are among the model’s stronger use cases, but they should not be treated as standardized benchmark results.

Tencent documents a high-reasoning default mode and a no-think setting available through the deployment chat template. High-reasoning mode is designed to spend more generation effort working through difficult problems. The no-think option can be useful when a faster, more direct response is preferable or when extended reasoning would add unnecessary latency and cost.

For coding, the model is positioned for code generation, debugging, repository-scale analysis, and coding agents that plan and execute multiple steps. Its long context can be useful for supplying large codebases, technical specifications, logs, and test output in one interaction. However, a large context window does not remove the need for testing. Generated code may still contain errors, misunderstand requirements, or make unsafe changes, especially when an agent has access to tools or files.

Tools, structured output, and deployment

Hy4 preview supports tool calling, which allows an application to describe functions that the model may request. The application, rather than the model itself, executes those functions and returns the results. This is useful for agents that need to search a database, manipulate files, run a build, call an internal service, or carry out a workflow.

Tencent also documents structured outputs. This can help applications request responses that follow a defined structure instead of relying on loosely formatted prose. Structured-output support should not automatically be interpreted as a separate, universally available JSON mode; the exact behavior depends on the serving interface and implementation being used.

The model supports streaming, so applications can receive generated text incrementally instead of waiting for the complete response. Tencent provides official deployment guidance for vLLM and SGLang, as well as an OpenAI-compatible local API. It also documents fine-tuning and cache support. These options make Hy4 preview relevant to organizations that want more control over hosting, integration, or model adaptation than a basic hosted chat interface provides.

Self-hosting a model with 770 billion total parameters is not a lightweight local-computer task. The open-weight release and deployment instructions improve control and flexibility, but they do not imply that ordinary consumer hardware can run the model efficiently. Infrastructure requirements, quantization choices, parallelism, and serving configuration will affect the actual cost and speed of deployment.

Pricing and speed trade-offs

Tencent Cloud TokenHub pricing is listed at CNY 6 per 1 million input tokens, CNY 0.30 per 1 million cached input tokens, and CNY 18 per 1 million output tokens. These are TokenHub prices and may vary by region, plan, or third-party access route. Cached-input pricing applies only when the provider’s caching conditions are met; it should not be assumed for every request.

Output is priced three times higher than uncached input on the listed TokenHub rates. Applications that repeatedly send the same long instructions or documents may benefit from caching if their integration qualifies, while applications that generate very long answers need to budget more carefully for output tokens.

The supplied editorial assessment gives Hy4 preview a speed score of 6 out of 10 and a cost score of 6 out of 10. These are comparative editorial judgments, not Tencent-published measurements. The model is unlikely to be the first choice for simple, high-volume requests where a smaller model can respond faster and more cheaply. Its value is more apparent when a task benefits from long context, extended reasoning, tool use, or complex coding behavior.

Main strengths and limitations

Strengths

  • Very large context: The 1-million-token context window is useful for large repositories, extensive documentation, long research collections, and multi-step project context.
  • Agent-oriented features: Tool calling, streaming, structured outputs, and an OpenAI-compatible API support practical application workflows.
  • Strong coding focus: Tencent positions the model for coding agents, automation, and complex technical work.
  • Deployment flexibility: Open weights, Apache 2.0 licensing, vLLM and SGLang guidance, and local API support offer more control than a closed chat-only service.
  • Reasoning controls: A high-reasoning default and a no-think option allow users to trade response depth against latency and cost.

Limitations

  • Preview status: Tencent identifies Hy4 preview as an early release. Production teams should expect possible behavior changes and should validate reliability before depending on it for critical workflows.
  • Text only: The model does not provide image, audio, or video input or output. A separate multimodal model is more appropriate for documents that require visual understanding, speech processing, or media generation.
  • Potentially lengthy reasoning: Tencent notes that the preview can produce lengthy reasoning and excessive self-verification on complex tasks. This can increase latency, token usage, and response length.
  • Infrastructure demands: The large total parameter count makes self-managed deployment substantially more demanding than deploying a small language model.
  • Unverified knowledge cutoff: No authoritative model-specific knowledge-cutoff date was identified in the supplied sources. Current information should therefore be provided through tools, retrieval, or application-managed data when freshness matters.

Best use cases for Hy4 preview

Hy4 preview is a good candidate for long-context coding agents that need to inspect many files, understand project conventions, propose changes, and work through tests or documentation. It can also suit software-development assistants that need tool calls rather than merely producing isolated code snippets.

Other suitable applications include document analysis across large text collections, productivity automation, technical research, scientific reasoning, and game development workflows. For example, an application could provide a large design specification, supporting documentation, and previous decisions in one context, then ask the model to identify inconsistencies or produce an implementation plan.

The model’s structured outputs and tool support are particularly useful when the response must feed another program. A workflow could ask Hy4 preview to classify incoming documents, return fields in a defined structure, and request a separate function when an external lookup is needed. The surrounding application must still validate both the structure and the substance of the result.

When to choose this model

Choose Hy4 preview when a project needs a very large text context, advanced coding or reasoning behavior, agent tools, and the flexibility of an open-weight model. It is especially compelling when the input consists of extensive text rather than media and when the additional latency or infrastructure cost is justified by the complexity of the task.

Choose a smaller or more mature language model when response speed, predictable production behavior, or low per-request cost matters more than maximum context and reasoning capacity. A smaller model may also be preferable for routine extraction, classification, short answers, and high-volume automation.

Choose a multimodal model instead when the core task involves images, scanned pages that require visual interpretation, audio, video, or media generation. Hy4 preview can process text descriptions of those materials, but the supplied specifications do not support direct media input or output.

Finally, consider a hosted alternative when operating large-scale inference infrastructure is impractical. Conversely, consider Hy4 preview’s open-weight and self-deployment options when data control, customization, or an Apache 2.0 license is more important than having the simplest managed service.

Bottom line

Hy4 preview is a technically ambitious Tencent language model built around long-context text processing, reasoning, coding, and tool-using agents. Its 1-million-token context, 64,000-token maximum output through TokenHub, open-weight availability, and deployment support give it a distinctive position for complex developer and research workflows.

The trade-off is that it remains a preview, can be demanding to run, and is not multimodal. Its pricing and speed are better justified for difficult tasks than for routine chat or inexpensive bulk processing. For users who can accept preview-stage limitations and need substantial text context with agent capabilities, Hy4 preview is a credible option; for simple, fast, media-oriented, or highly production-sensitive workloads, another model type may be a better fit.


Answers to Frequently Asked Questions

How can Hy4 preview be deployed and what license does it use?
Hy4 preview is available as an open-weight model under the Apache License 2.0. Tencent provides deployment guidance for vLLM and SGLang, along with an OpenAI-compatible local API. However, self-hosting is demanding because the model has 770 billion total parameters and may require substantial infrastructure, quantization, and parallelism.
Is Hy4 preview a multimodal model?
No. Hy4 preview is documented as a text-in, text-out model. It does not directly accept or generate images, audio, or video, so multimodal tasks require a separate model.
Can Hy4 preview be used for coding agents and tool calling?
Yes. Hy4 preview supports code generation, debugging, repository-scale analysis, coding agents, tool calling, streaming, structured outputs, and fine-tuning. Applications can use tool calls to let the model request actions such as database searches, file operations, builds, or internal API calls.
What is Hy4 preview?
Hy4 preview is Tencent’s open-weight, text-focused Mixture-of-Experts language model for long-context reasoning, coding, and tool-using agents. It has 770 billion total parameters and 49 billion active parameters per token, but it remains a preview release.
How large is Hy4 preview’s context window?
Hy4 preview has a headline context window of 1,000,000 tokens. Through Tencent Cloud TokenHub, the documented limits are up to 960,000 input tokens and 64,000 output tokens.


Sources 6
Provider

About Tencent AI