Hy-Image 3.5

Hy-Image-3.5-preview

by Tencent AI · Current preview model

Tencent Hy-Image-3.5-preview is a TokenHub preview model for high-resolution image generation and editing. It accepts text and reference images, supports multi-turn visual revisions and optional search enhancement, and produces synchronous results up to 4096×4096.

Image generation
Tencent Hy-Image-3.5-preview is designed for visual-production workflows rather than text conversation. The preview model combines text-to-image generation with reference-image conditioning and multi-turn editing, allowing users to refine an image over several requests while preserving conversational context. It supports 1K, 2K, and 4K output tiers through Tencent Cloud TokenHub.
Outputs

What Hy-Image-3.5-preview can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Web search Multimodal output
Model profile

Performance characteristics

7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Hy-Image 3.5
Model type Multimodal
Context window 100K tokens
Status Current preview model
Knowledge cutoff notes

Tencent's public API documentation specifies input limits and generation capabilities but does not publish a knowledge-cutoff date for this image-generation model.

Model notes

The canonical TokenHub request model ID is hy-image-v3.5-preview. The model uses an OpenAI Chat-style messages protocol and synchronously returns image results. It supports text-to-image, reference-image generation, multi-turn editing, multimodal text-plus-image input, external search enhancement, custom sizing, seeds, and watermark options. Documentation lists a 100,000-token input limit and width and height bounds of 256 to 8192 pixels, with a maximum area of 16,777,216 pixels and a practical maximum output of 4096×4096. The service performs internal SSE streaming but exposes a synchronous final response to callers. Editorial speed and cost scores are comparative estimates for image-generation APIs, not provider benchmarks.

Cost

Model pricing

Input 10 CNY per million tokens; reference-image generation uses the published TokenHub token rules
Output 15,000 tokens per 1K or 2K image; 20,000 tokens per 4K image, equivalent to approximately CNY 0.15 and CNY 0.20 respectively at the listed CNY 10 per million-token rate
Model guide

Tencent Hy-Image-3.5-preview: High-Resolution Image Generation with Multi-Turn Editing

Tencent Hy-Image-3.5-preview is a preview image-generation and editing model available through Tencent Cloud TokenHub. It accepts text and reference images through an OpenAI Chat-style messages API, supports multi-turn editing and optional external search enhancement, and can generate images at up to 4096×4096 resolution.

What is Hy-Image-3.5-preview?

Hy-Image-3.5-preview is Tencent's next-generation Hunyuan image-generation model for professional visual-content workflows. It is available as a preview service through Tencent Cloud TokenHub, where the canonical request model ID is hy-image-v3.5-preview.

The model creates images from natural-language prompts, but its main distinction is that it also accepts reference images and supports iterative editing. A request can combine text with one or more images, and later messages can use the earlier conversation as editing context. This makes the service more suitable for controlled visual development than a basic one-shot text-to-image endpoint.

Hy-Image-3.5-preview uses an OpenAI Chat-style messages protocol. That protocol is used to package text, reference images, and conversation history; it does not make the model a general-purpose text chatbot. Its direct output is an image result rather than a text completion.

Where it fits in Tencent's catalog

The model is positioned in Tencent Cloud TokenHub as a current preview image-generation service in the Hunyuan image family. Compared with the simpler Hy-Image-3.0 interface described in the supplied documentation, Hy-Image-3.5-preview uses a messages-based interface with multimodal inputs and multi-turn editing context.

That positioning matters when selecting an endpoint. The model is intended for users who need high-resolution image generation, image references, or an editing workflow. It is not presented as a replacement for a general language model, coding model, embedding model, or audio and video system.

Main capabilities

  • Text-to-image generation: creates an image from a natural-language description.
  • Reference-image generation: uses supplied images to condition or guide a new result.
  • Multi-turn editing: allows successive requests to refine an image using conversation history.
  • Multimodal prompts: combines written instructions with reference images.
  • Prompt and reasoning enhancement: the service can enhance prompts and perform internal processing before generation.
  • External search enhancement: the documented service supports an optional search-enhancement capability.
  • Generation controls: requests can include custom dimensions, automatic area-based sizing, seeds, and image-watermark options.

These features make the model useful when the desired result depends on more than a single textual description. For example, a user could provide a product reference image, request a promotional composition, and then ask for targeted revisions in later turns.

Input and output limits

The documentation lists a maximum input allowance of 100,000 tokens. For this image service, that limit is relevant to the complete request context, including text and conversation history, rather than representing a conventional text-model completion window. Long editing sessions should still be managed carefully because earlier messages and image references contribute to the request context.

Reference images can be supplied by URL or Base64 encoding. The international API documentation lists PNG, JPEG, and JPG support, a maximum file size of 20 MB per reference image, and support for up to 20 images in the documented request format.

Width and height can be specified within the documented 256-to-8192-pixel range. The maximum supported output area is 16,777,216 pixels, and the practical maximum advertised output is 4096×4096. The service therefore supports substantially larger image outputs than a workflow limited to small preview assets, although the requested dimensions remain subject to the model's area and tier rules.

Output is returned synchronously. In practical terms, the caller submits a request and receives the generated image result in the response rather than submitting a separate task and polling for completion. The documentation notes internal SSE processing, but user-facing token streaming is not exposed as a normal streaming text response.

Pricing and resolution tiers

TokenHub lists a reference price of 10 yuan per million tokens. Image generation is charged according to output-resolution tiers using the published token amounts:

Output tierPublished image chargeApproximate cost at the listed rate
1K15,000 tokens per imageApproximately CNY 0.15
2K15,000 tokens per imageApproximately CNY 0.15
4K20,000 tokens per imageApproximately CNY 0.20

The approximate currency figures are calculated from the published reference rate and should not be treated as a guaranteed final invoice. Actual billing depends on the TokenHub console, account configuration, and applicable service rules. The supplied research does not establish a separate recurring subscription price for this model.

The pricing structure creates a relatively clear quality-versus-cost trade-off: 1K and 2K generations use the same published token amount, while 4K output costs more. For drafts, social graphics, or early composition work, the lower tiers may be sufficient. The 4K tier is more appropriate when the image will be used in a high-resolution production workflow and the additional output area justifies the extra charge.

Supported modalities and model behavior

Hy-Image-3.5-preview supports text and image input, with image output. It does not provide audio or video input or output according to the supplied model record. Its multimodal capability therefore refers specifically to text-plus-image visual generation and editing.

The model record does not provide a reasoning score or coding score. It should not be evaluated as a reasoning or programming model: the service is built to interpret visual prompts and produce images, not to answer extended factual questions, write software, or complete arbitrary text. Similarly, no provider-published benchmark score is supplied for image quality, editing accuracy, or generation speed.

The optional external search enhancement is a service feature rather than evidence that the model is a general web-research assistant. It may help the image-generation workflow incorporate externally retrieved information, but the supplied documentation does not define it as a broad tool-calling or function-execution system. Tool-use support beyond that search-related capability is not verified.

Strengths and trade-offs

The strongest reason to use Hy-Image-3.5-preview is its combination of high-resolution output and iterative visual control. A simple text-to-image system may be adequate for generating independent concepts, but this model is better suited to workflows in which a user starts with a reference, evaluates the result, and requests several revisions.

Its 4K output option is useful for posters, product imagery, marketing graphics, and other assets that need more detail than a small preview. Support for custom sizing, seeds, and watermark options also gives production workflows more control than a minimal image endpoint.

The main trade-off is that the model is a preview service with a more specialized interface. Users who only need a quick, low-resolution image may not benefit from reference-image handling, long conversational context, or 4K generation. A simpler or faster image service could be more efficient for large batches of uncomplicated drafts, although the supplied research does not provide a direct latency benchmark for comparison.

Editorially, the model is rated as having a speed score of 7 and a cost score of 8 on a comparative internal scale. These are estimates for image-generation APIs, not Tencent benchmarks or published service-level guarantees. They should be used only as general orientation: the model offers low reference costs in the documented tiers, but actual response time and total expense depend on resolution, request size, account conditions, and workload.

Best use cases

  • Marketing assets: generate and revise campaign visuals, advertisements, and social-media graphics.
  • Product imagery: use a reference product image while exploring backgrounds, compositions, and presentation styles.
  • Posters and promotional artwork: take advantage of the 2K or 4K output tiers when detail and print-oriented resolution matter.
  • User-interface concepts: create visual directions, mockups, and interface-related imagery from detailed prompts.
  • Iterative image revisions: preserve conversational context while making targeted changes across multiple turns.
  • Reference-driven creative work: combine several supplied images with written instructions when a purely text-based prompt would be too ambiguous.

The model is especially appropriate when consistency across revisions matters. A user can describe the desired change in a later turn instead of rebuilding the entire prompt and re-uploading the full creative brief each time.

When to choose this model

Choose Hy-Image-3.5-preview when you need a Tencent Cloud image-generation endpoint with reference-image input, multi-turn editing, and output up to 4096×4096. It is a strong fit for visual teams that want to move from an initial concept to several controlled revisions within one conversational request flow.

Consider another type of option when the task is fundamentally different. A general-purpose language model is more appropriate for text analysis, coding, or conversational reasoning. A specialized video or audio model is necessary for those media types because Hy-Image-3.5-preview produces still images only. A simpler image generator may be preferable when you need only isolated drafts and do not require reference images, editing history, or high-resolution output.

Users should also consider a different service if exposed token streaming, asynchronous job management, or a guaranteed production release is essential. The supplied documentation describes synchronous final image responses and identifies this model as a preview service; it does not establish public streaming output or long-term availability guarantees.

Limitations to keep in mind

  • The model is a preview release, so its availability and behavior may change.
  • It is an image-generation and editing model, not a general text, coding, audio, video, or embedding model.
  • The documented maximum input allowance is 100,000 tokens, but image files, message history, and prompt content still need to fit within the service rules.
  • Reference images are subject to documented format, size, and quantity limits.
  • Although dimensions can be requested up to 8192 pixels on an individual width or height, the maximum area and practical advertised output limit must also be respected.
  • Published prices are reference rates; final billing depends on TokenHub account and service configuration.
  • No provider-published quality, reasoning, coding, or latency benchmarks are supplied in the available research.

Overall, Hy-Image-3.5-preview is best understood as a high-resolution, reference-aware image-generation service for iterative visual production. Its value comes less from general AI breadth than from combining multimodal prompts, multi-turn editing, and 4K-capable output in Tencent Cloud TokenHub.


Answers to Frequently Asked Questions

Who should use Hy-Image-3.5-preview?
It is best suited to visual teams creating marketing assets, product imagery, posters, promotional artwork, interface concepts, and other high-resolution content that requires reference images or repeated revisions. It is less suitable for text, coding, audio, video, or embedding tasks, and users needing guaranteed production availability should note that it is a preview service.
How much does image generation with Hy-Image-3.5-preview cost?
TokenHub lists a reference rate of 10 yuan per million tokens. The published image charges are 15,000 tokens for 1K and 2K output, or approximately CNY 0.15 each, and 20,000 tokens for 4K output, or approximately CNY 0.20. Actual billing depends on the TokenHub console, account configuration, and applicable service rules.
What image resolutions and input limits does Hy-Image-3.5-preview support?
The service supports requested widths and heights from 256 to 8192 pixels, with a maximum output area of 16,777,216 pixels and a practical advertised output of up to 4096×4096. Reference images can be PNG, JPEG, or JPG files up to 20 MB each, with up to 20 images in the documented request format. The maximum input allowance is 100,000 tokens.
What is Tencent Hy-Image-3.5-preview?
Tencent Hy-Image-3.5-preview is a preview image-generation and editing model available through Tencent Cloud TokenHub. It generates images from text, accepts reference images, and supports multi-turn editing through the model ID "hy-image-v3.5-preview".
Does Hy-Image-3.5-preview support reference images and multi-turn editing?
Yes. Hy-Image-3.5-preview accepts images supplied by URL or Base64 encoding and can combine them with text prompts. Later requests can use conversation history to refine or edit an earlier generated image.


Sources 6
Provider

About Tencent AI