WAND-Vega-Image-1.0-Pro is Tencent Cloud's professional image-generation model in the WAND-Vega Image 1.0 family. Its purpose is straightforward: turn a text description, one or more reference images, or a combination of both into a new image. The model is aimed at higher-quality visual production, with output options ranging from 1K to 4K resolution.
This makes it different from a conversational AI model. It does not generate text responses, write code, or provide general-purpose reasoning. Its value is in creating visual material such as brand imagery, product concepts, campaign assets, and other design-oriented images. The model is currently listed in Tencent Cloud TokenHub and uses an asynchronous task workflow: an application submits a generation request, receives a task_id, and checks the task until the image is ready.
What WAND-Vega-Image-1.0-Pro does
The model supports two central workflows:
- Text-to-image: a written prompt describes the desired subject, composition, style, lighting, or setting.
- Reference-to-image: one or more supplied images help guide the generated result. References can provide visual direction for a subject, product, composition, or overall look.
The Pro tier accepts prompts of up to 2,000 characters and supports up to six reference images according to Tencent Cloud's API documentation. The reference-image limit is especially relevant when a project needs to preserve several visual cues, although adding references can affect the request cost.
Available output sizes are 1K, 2K, and 4K. Tencent Cloud lists 1K and 2K generation at the same per-image price, while 4K output costs more. The available research does not establish the exact pixel dimensions represented by each resolution label, so the labels should be treated as the provider's supported output options rather than as a claim about a specific width and height.
Where the Pro model fits in Tencent Cloud's lineup
WAND-Vega-Image-1.0-Pro is part of Tencent's current WAND-Vega Image 1.0 catalog and is positioned as a quality-focused image-generation option. The canonical API model identifier is wand-vega-image-pro, while the catalog name is WAND-Vega-Image-1.0-Pro.
The Pro designation is useful for identifying the item within the WAND-Vega family, but the supplied documentation does not provide a full feature-by-feature comparison with other WAND-Vega variants. It is therefore more accurate to describe this model by its documented capabilities—reference-image input, multiple resolutions, and professional image-generation positioning—rather than to claim that it is universally better than every other model in the lineup.
Supported inputs and outputs
| Capability | Verified information |
|---|---|
| Text input | Supported, with prompts up to 2,000 characters |
| Image input | Supported through reference images |
| Maximum reference images | Up to six for the Pro tier |
| Image output | Supported at 1K, 2K, and 4K |
| Text output | Not the model's output type |
| Audio or video input/output | Not documented for this model |
| Generation workflow | Asynchronous task submission and polling |
The model's multimodal capability refers to its ability to accept both text and images as inputs and produce images as output. It should not be confused with a model that produces multiple media types. The documented output for WAND-Vega-Image-1.0-Pro is image generation.
Pricing for output and reference images
Tencent Cloud's listed Pro-tier pricing is:
| Item | Price |
|---|---|
| 1K image output | 0.95 CNY per image |
| 2K image output | 0.95 CNY per image |
| 4K image output | 1.71 CNY per image |
| First three reference images | Free |
| Each reference image after the first three | 0.10 CNY per image |
The reference-image rule applies to the images attached to a request. For example, a request with two references does not incur the additional reference-image charge, while a request with five references has two chargeable images at 0.10 CNY each. Output pricing is separate from this reference-image charge.
Tencent Cloud also describes the pricing in token terms, listing 10 CNY per million tokens. For practical planning, however, the per-image prices above are easier to use because they correspond directly to the documented 1K, 2K, and 4K output choices. The supplied pricing information does not provide a monthly subscription price or a guaranteed spending minimum.
API workflow and technical limits
WAND-Vega-Image-1.0-Pro is not documented as a synchronous image endpoint. Instead, the application submits a request and receives a task identifier. It then polls for completion and retrieves the resulting image when the task is finished. This design can work well for production queues and batch-style creative workflows, but it adds application logic and latency compared with an interface that returns an image immediately in the same request.
The documented input limit is a 2,000-character prompt. The research does not specify a context window, maximum output-token count, image file-size limit, or guaranteed completion time. Those values should not be inferred from the resolution options or from the model's asynchronous design.
The model is listed with streaming disabled. This is consistent with an image task that completes asynchronously rather than progressively streaming text tokens. The available information also does not verify fine-tuning, caching, batch API, function calling, or tool-use support. Developers should treat the documented image-generation task API as the supported integration surface and confirm any additional platform features in the current Tencent Cloud documentation.
Main strengths and trade-offs
The strongest documented advantages are visual input support, high-resolution output choices, and a relatively clear cost structure. Reference-to-image generation is useful when a prompt alone is not enough to communicate the desired appearance. A team can provide source imagery and use the prompt to describe the intended variation or final composition. The 4K option is also relevant when an image must serve as a larger design asset rather than only a small preview.
There are several practical limitations:
- It is specialized: the model is for image generation, not text generation, coding, embeddings, or conversational reasoning.
- Requests are asynchronous: an integration must manage task identifiers and polling instead of expecting a single immediate response.
- High resolution costs more: 4K generation is priced at 1.71 CNY per image, compared with 0.95 CNY for 1K or 2K.
- Extra references are chargeable: only the first three reference images are free; references four through six cost 0.10 CNY each.
- Some technical details are unspecified: the supplied documentation does not establish a completion-time guarantee, context window, or maximum file-size limit.
Editorially, the model is rated 7 out of 10 for speed and 7 out of 10 for cost in the supplied research. These are comparative editorial estimates, not Tencent Cloud benchmark results. They suggest a middle-ground assessment for image-generation workloads, but they should not be interpreted as a published latency or price ranking.
Reasoning, coding, and tool capabilities
WAND-Vega-Image-1.0-Pro does not have a documented reasoning score or a general-purpose reasoning mode. A prompt can describe a complex visual scene, but that should not be treated as evidence that the model can solve analytical problems or maintain a multi-turn argument.
Coding is likewise outside the model's stated purpose. It may be used inside a software application through Tencent Cloud's API, but that means the application is calling the model; it does not mean the model writes or executes code. The supplied research does not verify function calling, external tool use, web search, or structured JSON output. Developers who need those capabilities should use a model designed for language and tool-oriented tasks, while using WAND-Vega-Image-1.0-Pro as the image-generation component.
Best use cases
This model is a strong fit when the central deliverable is a generated image and visual quality matters more than conversational flexibility. Suitable examples include:
- Brand campaign concepts and advertising imagery
- Refined product images and alternate product presentations
- Creative direction boards based on one or more visual references
- High-resolution design assets for marketing and presentation work
- Image variations where a reference image needs to guide the generated result
A useful production pattern is to reserve 1K or 2K for exploration and iteration, then use 4K only for selected assets that need the higher output option. Keeping the number of reference images to three or fewer can also avoid the additional reference-image fee, provided that the smaller reference set gives the model enough direction.
When to choose WAND-Vega-Image-1.0-Pro
Choose WAND-Vega-Image-1.0-Pro when you need Tencent Cloud access to a dedicated image model that combines text prompts with reference images and offers a 4K output option. It is particularly appropriate for a workflow where image requests can run as asynchronous jobs and where the per-image pricing is easier to manage than a general subscription.
Another image model may be more appropriate if the project requires synchronous responses, documented real-time generation, a different style or editing workflow, or a lower-cost option at the required resolution. A language model is a better choice for writing copy, analyzing results, generating code, or coordinating tool calls. Likewise, a video or audio model is needed when the requested output is not a still image.
In short, WAND-Vega-Image-1.0-Pro is best evaluated as a specialized visual-production component. Its important decision points are the need for reference images, the desired resolution, the tolerance for asynchronous processing, and the cost of additional references—not general-purpose AI reasoning or coding performance.

