WAND-Vega-Image

WAND-Vega-Image-1.0-Flash

by Tencent AI · Current and available through Tencent Cloud TokenHub

Tencent Cloud's WAND-Vega-Image-1.0-Flash creates images from text prompts or reference images through an asynchronous TokenHub API. It supports prompts up to 2,000 characters, up to six PNG or JPEG references, 1K to 4K outputs, and reference costs of 0.45 to 1.008 yuan per image depending on resolution. Flash is positioned as a balanced alternative between the WAND-Vega Lite and Pro tiers.

Image generation Reasoning Coding
WAND-Vega-Image-1.0-Flash is Tencent Cloud's image-generation model for everyday creative production. It can create an image from a text description or use reference images to guide the result. Through Tencent Cloud TokenHub, developers can request 1K, 2K, or 4K images using an asynchronous API. The Flash tier is intended for users who need faster, more economical generation than a quality-focused model while retaining support for higher-resolution outputs.
Outputs

What WAND-Vega-Image-1.0-Flash can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family WAND-Vega-Image
Model type Multimodal
Status Current and available through Tencent Cloud TokenHub
Knowledge cutoff notes

A knowledge cutoff is not documented for this image-generation model. It is an image-generation endpoint rather than a general-purpose language model.

Model notes

The canonical API identifier is wand-vega-image-flash. The model supports text-to-image and reference-to-image generation with prompts up to 2,000 characters and output resolutions of 1K, 2K, and 4K. Flash accepts up to six PNG or JPEG reference images, each no larger than 10 MB, according to the API documentation. Image generation is asynchronous: submit a task, poll for completion, and retrieve a temporary image URL. The model is positioned between WAND-Vega-Image-1.0-Lite and WAND-Vega-Image-1.0-Pro. Editorial scores are comparative estimates for an image-generation model and do not represent vendor benchmarks.

Cost

Model pricing

Input 10 CNY per million tokens; input images are free for the Flash tier
Output Reference consumption: 45,000 tokens per 1K image (0.45 CNY), 67,500 tokens per 2K image (0.675 CNY), and 100,800 tokens per 4K image (1.008 CNY)
Model guide

WAND-Vega-Image-1.0-Flash for Fast, Cost-Conscious Image Generation

WAND-Vega-Image-1.0-Flash is Tencent Cloud's image-generation model for creating images from text prompts or reference images. Available through TokenHub, it supports 1K, 2K, and 4K output and is positioned between the Lite and Pro models in the WAND-Vega family, prioritizing a practical balance of speed, quality, and cost.

What is WAND-Vega-Image-1.0-Flash?

WAND-Vega-Image-1.0-Flash is an image-generation model provided by Tencent Cloud. Its primary job is to turn written prompts into images, although it can also use reference images to guide the composition or visual direction of a new image. It is not a general conversational model and does not produce text responses as its main output.

The model is available in Tencent Cloud TokenHub under the API identifier wand-vega-image-flash. TokenHub is the service layer through which developers submit image-generation tasks and retrieve their results. The model is designed for practical, repeated creative work rather than for text chat, coding assistance, or general reasoning.

Within the WAND-Vega family, Flash is positioned between WAND-Vega-Image-1.0-Lite and WAND-Vega-Image-1.0-Pro. Tencent Cloud describes the family structure in terms of trade-offs between speed, image quality, and cost. Flash is the middle option for users who need more than a basic low-cost tier but do not necessarily need the quality-focused Pro tier.

Inputs, outputs, and supported image capabilities

WAND-Vega-Image-1.0-Flash accepts text prompts and reference images. A prompt can contain up to 2,000 characters according to the API documentation. Reference images can be supplied as PNG or JPEG files, with each file limited to 10 MB. The Flash tier supports up to six reference images for a generation request.

Reference-to-image generation is useful when the request needs to preserve or draw inspiration from existing visual material. For example, a user might provide several product views, a character reference, or a visual style reference and then describe the desired scene in text. The supplied research confirms the input capability, but it does not document guarantees about exact identity preservation, pixel-level editing, or strict reproduction of every reference detail.

The model produces image output at three documented resolution levels:

  • 1K: the lowest supported output tier and the least expensive documented option.
  • 2K: a larger output suitable when additional detail is needed.
  • 4K: the highest supported output resolution and the most expensive of the listed options.

Supported aspect ratios include 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, and 21:9. This range covers square images, portrait and landscape formats, vertical social content, widescreen graphics, and wide cinematic layouts.

Where Flash fits in the WAND-Vega lineup

Flash is best understood as a balanced production tier. Lite is the lower-cost sibling, while Pro is positioned toward higher image quality and has different input-image charging rules. The available research does not provide a detailed quality benchmark or a complete feature-by-feature comparison between the three models, so the lineup positioning should not be read as a published benchmark ranking.

For a team generating many routine assets, Flash may be a more practical choice than a quality-maximizing model. It offers 4K output and reference-image input without requiring every request to use the Pro tier. Conversely, a user whose priority is the highest available quality within the WAND-Vega family may need to evaluate Pro instead. The appropriate choice depends on the visual quality required, the number of images being generated, and the acceptable cost per image.

How the TokenHub API workflow works

Image generation is asynchronous. Instead of keeping a single request open until the image is immediately returned, an application submits a generation task and receives a task identifier. It then polls the task-result endpoint until processing is complete. The completed result includes a temporary URL for the generated image.

The documented generation endpoint is:

https://tokenhub-intl.tencentcloudmaas.com/v1/wand/vega-images/generations

Requests use wand-vega-image-flash as the model value. In a production integration, the application should account for the delay between task submission and completion, handle unsuccessful task states, and download or persist the generated file promptly because the returned image URL is temporary.

The supplied documentation describes asynchronous task processing rather than a streaming response. Streaming is therefore not a documented feature for this model. The research also does not document tool calling, function calling, batch processing, fine-tuning, structured JSON output, or a conversational response mode.

Pricing and cost by resolution

Tencent Cloud lists Flash pricing at 10 yuan per million tokens. For this image-generation model, the published token figures can be translated into reference costs for each output resolution:

OutputReference token consumptionReference image cost
1K45,000 tokens0.45 yuan
2K67,500 tokens0.675 yuan
4K100,800 tokens1.008 yuan

These are reference prices calculated from the published token rate and consumption figures. Actual account billing, promotions, or other Tencent Cloud conditions may affect the final amount. Input images are listed as free for the Flash tier. The Pro tier follows a different input-image charging rule, so Pro pricing should not be applied to Flash requests.

The pricing structure creates a straightforward resolution trade-off: 1K is the lowest-cost choice for drafts and smaller placements, while 2K and 4K consume more tokens when the output needs additional size or detail. The research does not provide a separate monthly subscription price for the model.

Speed, quality, and practical trade-offs

Flash is intended to balance generation speed and image quality. This makes it a reasonable default for workflows that need a steady supply of usable images rather than a single carefully optimized hero image. Typical examples include social-media graphics, short-video covers, marketing materials, and other creative assets produced in batches or revised frequently.

Its main strength is the combination of practical cost, multiple output resolutions, and reference-image support. A creative team can begin with a lower-resolution generation, compare variations, and reserve 4K output for assets that need it. The 2,000-character prompt limit is also sufficient for detailed scene descriptions, although it does not make the model a general text-processing system.

The main limitation is that the available research does not establish a guaranteed quality level, benchmark score, or fixed generation time. Flash should therefore not be selected solely on the assumption that it will always be faster or visually better than every competing image model. Its speed and cost advantages are positioning claims and practical expectations, not independent benchmark results.

Reasoning, coding, and modality support

WAND-Vega-Image-1.0-Flash supports text input and image input, with image output. It does not produce text, audio, or video output. The model should be treated as a specialized image endpoint rather than a multimodal assistant that can discuss an image, write code, or execute tools in the same request.

Reasoning and coding are not meaningful primary capabilities for this model. The supplied editorial capability scores rate both reasoning and coding at 1 on a comparative scale, but these are editorial estimates, not Tencent-published benchmarks. The documentation also does not list tool use, function calling, structured output, or general-purpose reasoning features.

There is no documented context window or maximum text-output-token limit because the endpoint is designed for image generation rather than text completion. The relevant input constraints are the 2,000-character prompt limit, the supported reference-image formats and sizes, and the maximum of six reference images for Flash.

When to choose WAND-Vega-Image-1.0-Flash

Choose WAND-Vega-Image-1.0-Flash when the task is primarily image creation and the workflow needs a balance of cost, speed, resolution, and reference control. It is particularly suitable for:

  • Social-media posts and advertising graphics in portrait, square, or landscape formats.
  • Short-video covers and other recurring content assets.
  • Marketing concepts and product-presentation visuals.
  • Reference-guided variations based on a small collection of PNG or JPEG images.
  • Workflows that need occasional 4K output but do not require every image to use a premium tier.

Consider the Lite tier when the lowest possible generation cost is more important than the Flash balance of capabilities. Consider the Pro tier when the project prioritizes the higher-quality end of the WAND-Vega lineup and its different input-image pricing is acceptable. A conversational or coding model is more appropriate when the application needs text answers, programming help, reasoning, tool calls, or structured text responses rather than generated images.

Limitations to plan for

Before integrating Flash, account for the asynchronous request pattern and temporary result URLs. The application must poll for completion and store the image if it needs long-term access. It should also validate reference-image format and size before submission and avoid exceeding the six-image Flash limit.

Other limitations are documentation gaps rather than confirmed failures. The supplied material does not specify a knowledge cutoff, generation-time guarantee, benchmark performance, fine-tuning support, batch API, or detailed image-quality behavior. These capabilities should not be assumed. For a reliable deployment, test representative prompts and reference images using the intended resolutions and monitor the resulting cost and completion behavior.

Bottom line

WAND-Vega-Image-1.0-Flash is Tencent Cloud's middle-tier image-generation option for practical creative production. It combines text-to-image and reference-to-image generation, supports up to six reference images, offers 1K through 4K output, and uses an asynchronous TokenHub API. Its documented pricing is relatively easy to estimate by resolution, with reference costs of 0.45 yuan for 1K, 0.675 yuan for 2K, and 1.008 yuan for 4K output.

Flash is not a general-purpose AI assistant and should not be evaluated by conversational reasoning or coding standards. Its value lies in producing image assets at a balance of speed, cost, and quality. For routine graphics and reference-guided creation, it is a practical choice; for the lowest-cost generation or the highest-quality family option, the Lite and Pro tiers may be more appropriate.


Answers to Frequently Asked Questions

How does WAND-Vega-Image-1.0-Flash compare with the Lite and Pro tiers?
Flash is the middle option in the WAND-Vega family, positioned between the lower-cost Lite tier and the quality-focused Pro tier. It offers reference-image input and 4K output while balancing cost and production utility. Choose Lite when minimizing cost is the priority and consider Pro when the highest image quality is more important.
How does the WAND-Vega-Image-1.0-Flash API work?
The TokenHub API processes image generation asynchronously. An application submits a task to https://tokenhub-intl.tencentcloudmaas.com/v1/wand/vega-images/generations using the model value wand-vega-image-flash, receives a task identifier, and polls the task-result endpoint until processing is complete. The completed response includes a temporary image URL that should be downloaded or stored promptly.
How much does WAND-Vega-Image-1.0-Flash cost per generated image?
Tencent Cloud lists pricing at 10 yuan per million tokens. Based on the published token consumption, the reference cost is 0.45 yuan for 1K output, 0.675 yuan for 2K output, and 1.008 yuan for 4K output. Reference images are listed as free for the Flash tier, although actual billing may vary.
What is WAND-Vega-Image-1.0-Flash used for?
WAND-Vega-Image-1.0-Flash is a Tencent Cloud image-generation model for creating images from text prompts and reference images. It is suited to social-media graphics, marketing visuals, short-video covers, product presentations, and other recurring creative assets.
What image inputs and output formats does WAND-Vega-Image-1.0-Flash support?
The model accepts prompts of up to 2,000 characters and up to six PNG or JPEG reference images, with each image limited to 10 MB. It can generate images at 1K, 2K, or 4K resolution and supports aspect ratios including 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, and 21:9.


Sources 4
Provider

About Tencent AI