What is WAND-Vega-Image-1.0-Flash?
WAND-Vega-Image-1.0-Flash is an image-generation model provided by Tencent Cloud. Its primary job is to turn written prompts into images, although it can also use reference images to guide the composition or visual direction of a new image. It is not a general conversational model and does not produce text responses as its main output.
The model is available in Tencent Cloud TokenHub under the API identifier wand-vega-image-flash. TokenHub is the service layer through which developers submit image-generation tasks and retrieve their results. The model is designed for practical, repeated creative work rather than for text chat, coding assistance, or general reasoning.
Within the WAND-Vega family, Flash is positioned between WAND-Vega-Image-1.0-Lite and WAND-Vega-Image-1.0-Pro. Tencent Cloud describes the family structure in terms of trade-offs between speed, image quality, and cost. Flash is the middle option for users who need more than a basic low-cost tier but do not necessarily need the quality-focused Pro tier.
Inputs, outputs, and supported image capabilities
WAND-Vega-Image-1.0-Flash accepts text prompts and reference images. A prompt can contain up to 2,000 characters according to the API documentation. Reference images can be supplied as PNG or JPEG files, with each file limited to 10 MB. The Flash tier supports up to six reference images for a generation request.
Reference-to-image generation is useful when the request needs to preserve or draw inspiration from existing visual material. For example, a user might provide several product views, a character reference, or a visual style reference and then describe the desired scene in text. The supplied research confirms the input capability, but it does not document guarantees about exact identity preservation, pixel-level editing, or strict reproduction of every reference detail.
The model produces image output at three documented resolution levels:
- 1K: the lowest supported output tier and the least expensive documented option.
- 2K: a larger output suitable when additional detail is needed.
- 4K: the highest supported output resolution and the most expensive of the listed options.
Supported aspect ratios include 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, and 21:9. This range covers square images, portrait and landscape formats, vertical social content, widescreen graphics, and wide cinematic layouts.
Where Flash fits in the WAND-Vega lineup
Flash is best understood as a balanced production tier. Lite is the lower-cost sibling, while Pro is positioned toward higher image quality and has different input-image charging rules. The available research does not provide a detailed quality benchmark or a complete feature-by-feature comparison between the three models, so the lineup positioning should not be read as a published benchmark ranking.
For a team generating many routine assets, Flash may be a more practical choice than a quality-maximizing model. It offers 4K output and reference-image input without requiring every request to use the Pro tier. Conversely, a user whose priority is the highest available quality within the WAND-Vega family may need to evaluate Pro instead. The appropriate choice depends on the visual quality required, the number of images being generated, and the acceptable cost per image.
How the TokenHub API workflow works
Image generation is asynchronous. Instead of keeping a single request open until the image is immediately returned, an application submits a generation task and receives a task identifier. It then polls the task-result endpoint until processing is complete. The completed result includes a temporary URL for the generated image.
The documented generation endpoint is:
https://tokenhub-intl.tencentcloudmaas.com/v1/wand/vega-images/generations
Requests use wand-vega-image-flash as the model value. In a production integration, the application should account for the delay between task submission and completion, handle unsuccessful task states, and download or persist the generated file promptly because the returned image URL is temporary.
The supplied documentation describes asynchronous task processing rather than a streaming response. Streaming is therefore not a documented feature for this model. The research also does not document tool calling, function calling, batch processing, fine-tuning, structured JSON output, or a conversational response mode.
Pricing and cost by resolution
Tencent Cloud lists Flash pricing at 10 yuan per million tokens. For this image-generation model, the published token figures can be translated into reference costs for each output resolution:
| Output | Reference token consumption | Reference image cost |
|---|---|---|
| 1K | 45,000 tokens | 0.45 yuan |
| 2K | 67,500 tokens | 0.675 yuan |
| 4K | 100,800 tokens | 1.008 yuan |
These are reference prices calculated from the published token rate and consumption figures. Actual account billing, promotions, or other Tencent Cloud conditions may affect the final amount. Input images are listed as free for the Flash tier. The Pro tier follows a different input-image charging rule, so Pro pricing should not be applied to Flash requests.
The pricing structure creates a straightforward resolution trade-off: 1K is the lowest-cost choice for drafts and smaller placements, while 2K and 4K consume more tokens when the output needs additional size or detail. The research does not provide a separate monthly subscription price for the model.
Speed, quality, and practical trade-offs
Flash is intended to balance generation speed and image quality. This makes it a reasonable default for workflows that need a steady supply of usable images rather than a single carefully optimized hero image. Typical examples include social-media graphics, short-video covers, marketing materials, and other creative assets produced in batches or revised frequently.
Its main strength is the combination of practical cost, multiple output resolutions, and reference-image support. A creative team can begin with a lower-resolution generation, compare variations, and reserve 4K output for assets that need it. The 2,000-character prompt limit is also sufficient for detailed scene descriptions, although it does not make the model a general text-processing system.
The main limitation is that the available research does not establish a guaranteed quality level, benchmark score, or fixed generation time. Flash should therefore not be selected solely on the assumption that it will always be faster or visually better than every competing image model. Its speed and cost advantages are positioning claims and practical expectations, not independent benchmark results.
Reasoning, coding, and modality support
WAND-Vega-Image-1.0-Flash supports text input and image input, with image output. It does not produce text, audio, or video output. The model should be treated as a specialized image endpoint rather than a multimodal assistant that can discuss an image, write code, or execute tools in the same request.
Reasoning and coding are not meaningful primary capabilities for this model. The supplied editorial capability scores rate both reasoning and coding at 1 on a comparative scale, but these are editorial estimates, not Tencent-published benchmarks. The documentation also does not list tool use, function calling, structured output, or general-purpose reasoning features.
There is no documented context window or maximum text-output-token limit because the endpoint is designed for image generation rather than text completion. The relevant input constraints are the 2,000-character prompt limit, the supported reference-image formats and sizes, and the maximum of six reference images for Flash.
When to choose WAND-Vega-Image-1.0-Flash
Choose WAND-Vega-Image-1.0-Flash when the task is primarily image creation and the workflow needs a balance of cost, speed, resolution, and reference control. It is particularly suitable for:
- Social-media posts and advertising graphics in portrait, square, or landscape formats.
- Short-video covers and other recurring content assets.
- Marketing concepts and product-presentation visuals.
- Reference-guided variations based on a small collection of PNG or JPEG images.
- Workflows that need occasional 4K output but do not require every image to use a premium tier.
Consider the Lite tier when the lowest possible generation cost is more important than the Flash balance of capabilities. Consider the Pro tier when the project prioritizes the higher-quality end of the WAND-Vega lineup and its different input-image pricing is acceptable. A conversational or coding model is more appropriate when the application needs text answers, programming help, reasoning, tool calls, or structured text responses rather than generated images.
Limitations to plan for
Before integrating Flash, account for the asynchronous request pattern and temporary result URLs. The application must poll for completion and store the image if it needs long-term access. It should also validate reference-image format and size before submission and avoid exceeding the six-image Flash limit.
Other limitations are documentation gaps rather than confirmed failures. The supplied material does not specify a knowledge cutoff, generation-time guarantee, benchmark performance, fine-tuning support, batch API, or detailed image-quality behavior. These capabilities should not be assumed. For a reliable deployment, test representative prompts and reference images using the intended resolutions and monitor the resulting cost and completion behavior.
Bottom line
WAND-Vega-Image-1.0-Flash is Tencent Cloud's middle-tier image-generation option for practical creative production. It combines text-to-image and reference-to-image generation, supports up to six reference images, offers 1K through 4K output, and uses an asynchronous TokenHub API. Its documented pricing is relatively easy to estimate by resolution, with reference costs of 0.45 yuan for 1K, 0.675 yuan for 2K, and 1.008 yuan for 4K output.
Flash is not a general-purpose AI assistant and should not be evaluated by conversational reasoning or coding standards. Its value lies in producing image assets at a balance of speed, cost, and quality. For routine graphics and reference-guided creation, it is a practical choice; for the lowest-cost generation or the highest-quality family option, the Lite and Pro tiers may be more appropriate.

