What WAND-Vega-Video-1.0-Lite is
WAND-Vega-Video-1.0-Lite is a lightweight video-generation model provided by Tencent Cloud and listed in the company’s TokenHub catalog. Its main role is to turn a text description and, when needed, reference media into a short generated video. The “Lite” positioning is relevant for teams that value production speed and controlled costs over long clips or the broadest possible generation feature set.
The model is aimed at practical content production: product demonstrations, advertising variations, social-media clips, image-to-video animation, and other workflows that create many short assets. It is not documented as a general-purpose language model, so text reasoning, coding, tool use, and conventional text output are not its purpose.
Tencent Cloud currently documents three main generation patterns: text-to-video, first-and-last-frame generation, and reference-based video generation. Every request must include at least one text element, even when reference media are also supplied.
Three ways to guide a generation
Text-to-video
In the simplest workflow, the user supplies a text prompt describing the desired scene, subject, action, style, or camera movement. The model then generates a short video based on that instruction. This is suitable for creating a concept clip, a simple social post, or an initial product-video draft without supplying source media.
First-and-last-frame generation
The model can use up to one first-frame image and one last-frame image to guide how a video begins and ends. This is useful when the transition itself matters—for example, showing a product changing position, moving from one scene to another, or transforming between two visual states.
The two images act as visual anchors rather than merely decorative references. The generated frames between them are still produced by the model, so the feature should be treated as guided transition generation rather than a guarantee of exact object or motion continuity.
Reference-based video generation
Reference-based generation allows a prompt to be combined with reference images, videos, and audio. These inputs can provide guidance about appearance, movement, characters, scenes, or sound. Tencent Cloud recommends using at least one reference image or video for this mode and keeping the total reference assets within the documented limits.
This makes the Lite model more suitable than a text-only workflow when a campaign needs to preserve a product appearance, draw on an existing motion sample, or coordinate several kinds of creative direction. The supplied research does not specify a fixed maximum number of assets, so implementations should follow the limits in the current Tencent Cloud API documentation rather than assume an undocumented allowance.
Supported specifications
| Specification | Documented support |
|---|---|
| Model ID | wand-vega-video-1.0-lite |
| Video duration | 4 to 15 seconds; 5 seconds by default |
| Resolution | 768P, 1080P, 2K, and 4K |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and adaptive |
| Inputs | Text, image URLs, video URLs, and audio URLs |
| Output | Generated video URL, with an optional last-frame image |
| Processing | Asynchronous task submission and polling |
| Default concurrency | Five, according to the listed model information |
The available resolution range is unusually broad for a lightweight short-video model, extending from 768P through 4K. However, supported values are discrete: a request outside the documented range is rejected rather than automatically downgraded. Teams should validate resolution and aspect-ratio choices before submitting production jobs.
How the TokenHub workflow operates
WAND-Vega-Video-1.0-Lite is accessed through Tencent Cloud TokenHub using a task-based video endpoint. An application submits the model ID, a content array containing the prompt and any reference media, and optional generation settings such as duration, resolution, aspect ratio, and output options.
The service returns a task identifier rather than an immediately available video. The application then polls the task until it succeeds or fails. This asynchronous design fits batch production, where a queue can manage many jobs, but it is less convenient for an interface that expects a video to appear immediately.
Successful video URLs are temporary and are documented as remaining valid for 12 hours. Production systems should download each result and move it to durable storage promptly. A robust integration should also handle queued tasks, failures, retries, output transfer, and expiration of unused links.
Strengths and trade-offs
The model’s clearest strength is the combination of short-video generation and multiple reference modalities. A team can start with a text-only concept, add opening and closing frames for a controlled transition, or use images, video, and audio to provide richer direction. The range of durations, aspect ratios, and resolutions also covers common advertising, e-commerce, and social-media formats.
Its lightweight positioning is another practical advantage for high-volume workflows. The supplied editorial assessment rates its speed and cost characteristics at 8 out of 10. Those scores are editorial evaluations, not Tencent Cloud benchmarks or provider-published guarantees. They indicate that the model appears well suited to fast, cost-conscious production, while the exact cost and latency for a workload still depend on service conditions and request settings.
The trade-off is scope. The model generates short clips rather than long-form video, and it does not provide a documented streaming-generation interface. Its asynchronous operation adds queue management and polling work. Higher resolution, longer duration, or more complex reference inputs may also affect practical throughput, but the supplied documentation does not provide verified latency or per-generation performance figures.
Pricing and limits
Tencent Cloud’s current public pricing information does not expose a verified model-specific price for WAND-Vega-Video-1.0-Lite. There is therefore no reliable price-per-video figure to report. Prospective users should confirm the applicable TokenHub billing terms directly before estimating campaign costs.
The published material also does not specify a knowledge cutoff, text context window, maximum text-token output, fine-tuning support, batch API, or a separate release date for this exact model. These omissions are important when evaluating the model: they mean those capabilities should be treated as unknown, not as supported by default.
Some limits are explicit. The generated duration must be between 4 and 15 seconds, supported resolution values must be selected, and a text element is required. The video output link is temporary. Other operational limits, including the precise maximum number of reference assets, should be checked against the current API guide.
Modalities and capability profile
WAND-Vega-Video-1.0-Lite accepts text, image, video, and audio inputs and produces video output. It can optionally return a last-frame image alongside the generated video. In practical terms, it is a multimodal video model: the media inputs influence a visual or audiovisual generation task rather than being used for open-ended conversation.
| Capability | Assessment |
|---|---|
| Text input | Supported and required |
| Image input | Supported through image references and frame controls |
| Video input | Supported as a reference input |
| Audio input | Supported as a reference input |
| Video output | Supported |
| Text output | Not the model’s documented output type |
| Reasoning or coding | Not documented as model capabilities |
| Tool or function calling | Not documented |
| Streaming output | Not supported; generation is asynchronous |
These distinctions matter when selecting the model for an application. WAND-Vega-Video-1.0-Lite should be paired with separate software or models for prompt construction, workflow control, metadata generation, content moderation, or other text-based tasks unless those functions are implemented by the surrounding application.
Best use cases
- E-commerce product clips: Generate multiple short demonstrations or promotional variations from product imagery and a text brief.
- Social and advertising assets: Produce vertical, square, widescreen, or other aspect-ratio variants for short-form campaigns.
- Image-to-video animation: Turn a still product, character, or scene into a moving clip.
- First-to-last-frame transitions: Create a controlled visual change between two supplied states.
- Reference-guided production: Combine images, sample video, and audio with a prompt when a text-only instruction is not specific enough.
- Batch generation: Submit many asynchronous tasks for later review and storage.
When to choose this model
Choose WAND-Vega-Video-1.0-Lite when the requirement is a short generated clip, especially when production volume, multiple reference types, and cost awareness matter more than real-time response or long-duration continuity. It is a reasonable fit for teams building a queued content pipeline rather than a live interactive video-generation experience.
Another option may be more appropriate when the project needs long-form video, documented streaming generation, fine-tuning, a batch API with formally specified behavior, or a published model-specific price that can be used directly for budgeting. A different model may also be preferable when the workflow requires text reasoning, coding, tool calls, or structured text output in the same model. Those requirements are not documented for WAND-Vega-Video-1.0-Lite.
For evaluation, begin with representative prompts and reference media from the intended campaign. Test the actual duration, resolution, aspect-ratio, polling, retry, and storage workflow together; a model that produces suitable clips is only operationally useful if the surrounding system can preserve its temporary outputs and handle asynchronous completion reliably.

