What is Hy4 preview?
Hy4 preview is a text-focused large language model from Tencent. Tencent describes it as a flagship Mixture-of-Experts, or MoE, model. In an MoE system, the model contains many specialist neural-network components, but only a selected subset is activated for each token. This allows a model to have a very large overall capacity without using every parameter for every piece of text.
Hy4 preview has 770 billion total parameters and 49 billion active parameters per token. Those figures describe the model’s architecture, not a guarantee of benchmark performance. In practical terms, the design is intended to provide substantial reasoning and language capacity while reducing the amount of computation used for each generated token compared with a dense model of the same total size.
The model was released as a preview and is available as an open-weight model, through Tencent products and Tencent Cloud TokenHub, and through third-party gateways including OpenRouter according to the supplied research. Tencent also provides documentation for local or self-managed deployment. Because it is a preview, users should treat its behavior, interfaces, availability, and operational reliability as less settled than those of a mature production release.
Where Hy4 preview fits in Tencent’s lineup
Hy4 preview belongs to Tencent’s Hunyuan model family and represents a high-end general-purpose language model in that family. It is distinct from Yuanbao, Tencent’s consumer AI assistant, even though Tencent’s model generations may be integrated into Yuanbao and other Tencent products. Hy4 preview should therefore be evaluated as a model that can be deployed or accessed through developer-oriented channels, not as a consumer chat subscription.
Its position is also different from Tencent’s image, video, 3D, and other multimodal services. Hy4 preview is documented as a text-in, text-out model. It does not accept images, audio, or video as model inputs, and it does not directly generate those media types. That makes it a focused language and agent model rather than a single model for every modality in Tencent’s broader AI ecosystem.
Key specifications and limits
| Specification | Hy4 preview |
|---|---|
| Provider | Tencent |
| Model family | Hy4 |
| Status | Preview |
| Architecture | Mixture-of-Experts |
| Total parameters | 770 billion |
| Active parameters per token | 49 billion |
| Context window | 1,000,000 tokens |
| Maximum input through TokenHub | 960,000 tokens |
| Maximum output through TokenHub | 64,000 tokens |
| Input and output modalities | Text input and text output |
| License | Apache License 2.0 |
| Tool use | Supported |
| Streaming | Supported |
| Fine-tuning | Supported |
| Structured outputs | Supported |
The advertised 1-million-token context window is the model’s headline capacity. Tencent’s TokenHub documentation specifies up to 960,000 input tokens and 64,000 output tokens in that overall context arrangement. A token is a small unit of text used by language models; the exact number of words represented by a token count varies by language and content. The practical implication is that Hy4 preview can work with unusually large collections of text, although context capacity alone does not guarantee that every detail will be handled perfectly.
Reasoning and coding capabilities
Hy4 preview is intended for tasks that require more than short-form text completion. The supplied research rates its reasoning and coding capabilities at 8 out of 10 as editorial evaluations, not as scores published by Tencent. Those ratings indicate an assessment that reasoning and programming are among the model’s stronger use cases, but they should not be treated as standardized benchmark results.
Tencent documents a high-reasoning default mode and a no-think setting available through the deployment chat template. High-reasoning mode is designed to spend more generation effort working through difficult problems. The no-think option can be useful when a faster, more direct response is preferable or when extended reasoning would add unnecessary latency and cost.
For coding, the model is positioned for code generation, debugging, repository-scale analysis, and coding agents that plan and execute multiple steps. Its long context can be useful for supplying large codebases, technical specifications, logs, and test output in one interaction. However, a large context window does not remove the need for testing. Generated code may still contain errors, misunderstand requirements, or make unsafe changes, especially when an agent has access to tools or files.
Tools, structured output, and deployment
Hy4 preview supports tool calling, which allows an application to describe functions that the model may request. The application, rather than the model itself, executes those functions and returns the results. This is useful for agents that need to search a database, manipulate files, run a build, call an internal service, or carry out a workflow.
Tencent also documents structured outputs. This can help applications request responses that follow a defined structure instead of relying on loosely formatted prose. Structured-output support should not automatically be interpreted as a separate, universally available JSON mode; the exact behavior depends on the serving interface and implementation being used.
The model supports streaming, so applications can receive generated text incrementally instead of waiting for the complete response. Tencent provides official deployment guidance for vLLM and SGLang, as well as an OpenAI-compatible local API. It also documents fine-tuning and cache support. These options make Hy4 preview relevant to organizations that want more control over hosting, integration, or model adaptation than a basic hosted chat interface provides.
Self-hosting a model with 770 billion total parameters is not a lightweight local-computer task. The open-weight release and deployment instructions improve control and flexibility, but they do not imply that ordinary consumer hardware can run the model efficiently. Infrastructure requirements, quantization choices, parallelism, and serving configuration will affect the actual cost and speed of deployment.
Pricing and speed trade-offs
Tencent Cloud TokenHub pricing is listed at CNY 6 per 1 million input tokens, CNY 0.30 per 1 million cached input tokens, and CNY 18 per 1 million output tokens. These are TokenHub prices and may vary by region, plan, or third-party access route. Cached-input pricing applies only when the provider’s caching conditions are met; it should not be assumed for every request.
Output is priced three times higher than uncached input on the listed TokenHub rates. Applications that repeatedly send the same long instructions or documents may benefit from caching if their integration qualifies, while applications that generate very long answers need to budget more carefully for output tokens.
The supplied editorial assessment gives Hy4 preview a speed score of 6 out of 10 and a cost score of 6 out of 10. These are comparative editorial judgments, not Tencent-published measurements. The model is unlikely to be the first choice for simple, high-volume requests where a smaller model can respond faster and more cheaply. Its value is more apparent when a task benefits from long context, extended reasoning, tool use, or complex coding behavior.
Main strengths and limitations
Strengths
- Very large context: The 1-million-token context window is useful for large repositories, extensive documentation, long research collections, and multi-step project context.
- Agent-oriented features: Tool calling, streaming, structured outputs, and an OpenAI-compatible API support practical application workflows.
- Strong coding focus: Tencent positions the model for coding agents, automation, and complex technical work.
- Deployment flexibility: Open weights, Apache 2.0 licensing, vLLM and SGLang guidance, and local API support offer more control than a closed chat-only service.
- Reasoning controls: A high-reasoning default and a no-think option allow users to trade response depth against latency and cost.
Limitations
- Preview status: Tencent identifies Hy4 preview as an early release. Production teams should expect possible behavior changes and should validate reliability before depending on it for critical workflows.
- Text only: The model does not provide image, audio, or video input or output. A separate multimodal model is more appropriate for documents that require visual understanding, speech processing, or media generation.
- Potentially lengthy reasoning: Tencent notes that the preview can produce lengthy reasoning and excessive self-verification on complex tasks. This can increase latency, token usage, and response length.
- Infrastructure demands: The large total parameter count makes self-managed deployment substantially more demanding than deploying a small language model.
- Unverified knowledge cutoff: No authoritative model-specific knowledge-cutoff date was identified in the supplied sources. Current information should therefore be provided through tools, retrieval, or application-managed data when freshness matters.
Best use cases for Hy4 preview
Hy4 preview is a good candidate for long-context coding agents that need to inspect many files, understand project conventions, propose changes, and work through tests or documentation. It can also suit software-development assistants that need tool calls rather than merely producing isolated code snippets.
Other suitable applications include document analysis across large text collections, productivity automation, technical research, scientific reasoning, and game development workflows. For example, an application could provide a large design specification, supporting documentation, and previous decisions in one context, then ask the model to identify inconsistencies or produce an implementation plan.
The model’s structured outputs and tool support are particularly useful when the response must feed another program. A workflow could ask Hy4 preview to classify incoming documents, return fields in a defined structure, and request a separate function when an external lookup is needed. The surrounding application must still validate both the structure and the substance of the result.
When to choose this model
Choose Hy4 preview when a project needs a very large text context, advanced coding or reasoning behavior, agent tools, and the flexibility of an open-weight model. It is especially compelling when the input consists of extensive text rather than media and when the additional latency or infrastructure cost is justified by the complexity of the task.
Choose a smaller or more mature language model when response speed, predictable production behavior, or low per-request cost matters more than maximum context and reasoning capacity. A smaller model may also be preferable for routine extraction, classification, short answers, and high-volume automation.
Choose a multimodal model instead when the core task involves images, scanned pages that require visual interpretation, audio, video, or media generation. Hy4 preview can process text descriptions of those materials, but the supplied specifications do not support direct media input or output.
Finally, consider a hosted alternative when operating large-scale inference infrastructure is impractical. Conversely, consider Hy4 preview’s open-weight and self-deployment options when data control, customization, or an Apache 2.0 license is more important than having the simplest managed service.
Bottom line
Hy4 preview is a technically ambitious Tencent language model built around long-context text processing, reasoning, coding, and tool-using agents. Its 1-million-token context, 64,000-token maximum output through TokenHub, open-weight availability, and deployment support give it a distinctive position for complex developer and research workflows.
The trade-off is that it remains a preview, can be demanding to run, and is not multimodal. Its pricing and speed are better justified for difficult tasks than for routine chat or inexpensive bulk processing. For users who can accept preview-stage limitations and need substantial text context with agent capabilities, Hy4 preview is a credible option; for simple, fast, media-oriented, or highly production-sensitive workloads, another model type may be a better fit.

