What Hy-Image-3.0 is designed to do
Hy-Image-3.0 is Tencent’s Hunyuan image-generation model. It accepts a text description and returns a generated image, and it can also use one or more reference images to guide the result. This makes it suitable for workflows where a prompt alone is not enough—for example, preserving the visual direction of a product image, using several creative references, or adapting an existing concept to a different composition.
The managed API exposes the model through Tencent Cloud TokenHub with the identifier hy-image-v3. Tencent also publishes an open-weight project called HunyuanImage-3.0. These are related release forms of the same current model family, but they serve different audiences: TokenHub is the simpler option for applications that need a hosted service, while the open-weight release is intended for organizations able to operate substantial GPU infrastructure.
Tencent describes the model as a native multimodal image-generation system that combines multimodal understanding with image generation. In practical terms, its supported workflow is image creation rather than general conversation, document analysis, or language-model completion.
Inputs, outputs, and core capabilities
Hy-Image-3.0 supports two main generation modes:
- Text-to-image: create an image from a written prompt.
- Reference-image generation: create an image using one or more supplied images as visual guidance.
Reference images can be provided through URLs or Base64 data. The documented API accepts PNG, JPEG, and JPG images, with up to three reference images and a maximum size of 10 MB per image. The output is a generated image; the supplied research does not identify a text, audio, video, embedding, or speech output mode for this model.
Additional controls make the service more suitable for production image workflows than a minimal prompt-only endpoint:
- Automatic prompt rewriting can expand or optimize a user’s prompt before generation.
- Custom width and height settings support different compositions and aspect ratios.
- Preset sizes cover square, wide, tall, and panoramic formats.
- A seed can make single-image generation reproducible.
- Optional watermark footnotes are available.
If the size parameter is omitted, the service can select an appropriate output shape automatically. This is convenient for basic applications, while explicit dimensions are more useful when the generated image must fit a known banner, product-card, or publishing layout.
Image sizes, seeds, and processing controls
The API documentation specifies that explicit width and height values must fall between 512 and 2048 pixels. It also describes an area limit of no more than 1024 × 1024 pixels for the standard size parameter, so developers should check the applicable Tencent Cloud documentation when selecting dimensions. Listed presets include 1024 × 1024, 1280 × 720, 768 × 1280, and 2048 × 512.
Hy-Image-3.0 supports seed values from 1 through 4,294,967,295. When one image is generated, supplying the same seed can help reproduce a result. A missing seed or a seed of zero causes the service to use a random seed. Reproducibility should therefore be understood as a generation control, not a guarantee that every broader application workflow will remain identical after service changes.
Prompt rewriting is another optional control. It can improve or expand a short user instruction, but Tencent’s documentation notes that enabling it may add approximately 11 seconds to processing time. Applications that prioritize response speed may prefer to rewrite prompts themselves or disable the service-side option when the input is already carefully structured.
Position in Tencent’s Hunyuan catalog
Hy-Image-3.0 belongs to Tencent’s Hunyuan family but is specialized for image generation. It should not be evaluated as a general-purpose language model: the available record lists no context window, maximum text-output limit, reasoning score, coding capability, tool-use support, or structured-output mode for it. Its relevant input modalities are text and images, and its primary output modality is images.
The open-weight HunyuanImage-3.0 repository includes a base model, instruct variants, and distilled variants. The base model is primarily intended for text-to-image generation. The instruct variants add reference-image generation, prompt self-rewriting, and chain-of-thought-oriented generation workflows. These variants are distinct from the managed record described here, so their capabilities and deployment requirements should not automatically be treated as properties of every TokenHub configuration.
TokenHub API versus open-weight deployment
TokenHub is the practical route for developers who want to integrate Hy-Image-3.0 without operating the model themselves. The service handles inference infrastructure and exposes synchronous image generation through Tencent Cloud. Billing is based on image-generation usage or the applicable regional token-equivalent pricing rather than conventional language-model input and output tokens.
The open-weight release provides more control but has demanding hardware requirements. Tencent’s model materials list approximately 80 billion total parameters and 13 billion active parameters for the base HunyuanImage-3.0 model. The model card recommends at least three 80 GB GPUs for the base version and at least eight 80 GB GPUs for the instruct variants. These requirements apply to self-managed inference and do not directly describe the resources required by the hosted TokenHub service.
For most teams, the distinction is straightforward: choose TokenHub when deployment simplicity, managed scaling, and an API integration matter most; consider the open-weight release only when control over deployment and model operation justifies the infrastructure investment.
Pricing and cost considerations
Pricing varies by Tencent Cloud region and billing surface. The cited mainland China TokenHub pricing lists a reference price of approximately 0.20 CNY per generated image. Tencent’s cited international pricing documentation lists approximately 0.032 USD per image, with image generation described as 20,000 tokens per image at 1.6 USD per million tokens.
These figures are regional reference prices rather than a universal global rate. The applicable Tencent Cloud console and pricing page should be checked before budgeting, particularly because taxes, account region, promotions, and service changes can affect the final charge.
From a practical cost perspective, Hy-Image-3.0 is more naturally compared with other hosted image-generation services than with text models. Its usage cost is tied to image creation, and prompt rewriting can increase latency even when it does not change the basic image-generation purpose. The model may be a good fit when reference control and flexible composition are more valuable than the absolute lowest-cost or fastest image endpoint.
Strengths and limitations
Its principal strengths are the combination of text and reference-image generation, support for up to three references, custom output dimensions, and production-oriented controls such as seeds and optional prompt rewriting. These features can reduce the need for separate preprocessing steps in applications that create marketing graphics, product visuals, concept art, or other controlled image variations.
There are also important limitations:
- It is an image-generation model, not a general-purpose assistant, coding model, embedding service, or conversational endpoint.
- The supplied research does not specify a conventional context window or maximum output-token limit because the model does not produce language-model text output.
- Reference inputs are limited to three images, and each image is limited to 10 MB according to the API documentation.
- Explicit dimensions are subject to documented width, height, and area constraints.
- Prompt rewriting may add approximately 11 seconds of processing time.
- Local operation requires substantial multi-GPU hardware, especially for instruct variants.
- The managed service is synchronous, and the supplied record does not identify streaming, tool calling, function calling, or batch API support.
The model’s documented capabilities also do not establish general image understanding as a separate analysis endpoint. Its image inputs are described in the context of guiding image generation.
When to choose Hy-Image-3.0
Hy-Image-3.0 is a strong candidate when an application needs hosted image creation with more control than a basic text-to-image prompt. Suitable examples include:
- Marketing and advertising creative generation.
- E-commerce product imagery and variations.
- Concept development and visual ideation.
- Applications that combine a user prompt with several visual references.
- Creative tools that need wide, tall, square, or panoramic output formats.
- Workflows where repeatable single-image generation through a seed is useful.
Another option may be more appropriate when the primary requirement is text generation, software development, speech, video, embeddings, or multimodal question answering. A hosted image service with lower latency may also be preferable when prompt rewriting is unnecessary and response time is more important than its additional optimization step. Conversely, the open-weight release is worth considering only for teams that need self-managed deployment and can provide the recommended GPU capacity.
Bottom line
Hy-Image-3.0 is a specialized Tencent Hunyuan image-generation model whose most useful distinction is reference-guided creation alongside ordinary text-to-image generation. TokenHub provides the accessible managed route, while the open-weight HunyuanImage-3.0 release targets infrastructure-capable users. Its support for three reference images, custom dimensions, prompt rewriting, and seed control makes it relevant to structured creative workflows, but it should be selected as an image-generation system—not as a replacement for a general language or multimodal assistant.

