WAND-Vega-Video

WAND-Vega-Video-1.0-Lite

by Tencent AI · Current and accessible through Tencent Cloud TokenHub

Tencent Cloud’s WAND-Vega-Video-1.0-Lite is a lightweight asynchronous video model for text-to-video, first-and-last-frame transitions, and reference-guided generation. It accepts text, images, videos, and audio, producing 4–15 second clips at 768P through 4K. It is best suited to cost-conscious batch production, while pricing, context limits, fine-tuning, and several advanced capabilities remain undocumented.

Video generation
WAND-Vega-Video-1.0-Lite is a current Tencent Cloud TokenHub model designed for fast, lower-cost video production rather than long-form filmmaking or real-time generation. It can create short clips from text, guide a transition between a first and last frame, or use reference images, videos, and audio to influence the result.
Outputs

What WAND-Vega-Video-1.0-Lite can produce

Video generation
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family WAND-Vega-Video
Model type Lightweight
Status Current and accessible through Tencent Cloud TokenHub
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for this video-generation model. Its documented generation behavior is controlled by prompts and reference media rather than a stated textual knowledge cutoff.

Model notes

Canonical model ID: wand-vega-video-1.0-lite. The current TokenHub catalog lists the model with text-to-video, first-and-last-frame generation, and reference-based video generation. The API accepts text, image, video, and audio reference inputs. Supported output durations are 4–15 seconds, with 768P, 1080P, 2K, and 4K resolution options. Video generation is asynchronous; generated video URLs are temporary and documented as valid for 12 hours. Tencent Cloud’s current public pricing page does not expose a verified model-specific price for this exact video model. The published documentation does not provide a knowledge cutoff, context length, maximum output-token limit, or separate model release date.

Model guide

WAND-Vega-Video-1.0-Lite: Cost-Efficient Short Video Generation with Multimodal References

WAND-Vega-Video-1.0-Lite is Tencent Cloud’s lightweight video-generation model for short, cost-conscious production workflows. It supports text-to-video, first-and-last-frame transitions, and reference-guided generation using text, images, videos, and audio. The model produces 4- to 15-second videos at resolutions from 768P to 4K through an asynchronous TokenHub API.

What WAND-Vega-Video-1.0-Lite is

WAND-Vega-Video-1.0-Lite is a lightweight video-generation model provided by Tencent Cloud and listed in the company’s TokenHub catalog. Its main role is to turn a text description and, when needed, reference media into a short generated video. The “Lite” positioning is relevant for teams that value production speed and controlled costs over long clips or the broadest possible generation feature set.

The model is aimed at practical content production: product demonstrations, advertising variations, social-media clips, image-to-video animation, and other workflows that create many short assets. It is not documented as a general-purpose language model, so text reasoning, coding, tool use, and conventional text output are not its purpose.

Tencent Cloud currently documents three main generation patterns: text-to-video, first-and-last-frame generation, and reference-based video generation. Every request must include at least one text element, even when reference media are also supplied.

Three ways to guide a generation

Text-to-video

In the simplest workflow, the user supplies a text prompt describing the desired scene, subject, action, style, or camera movement. The model then generates a short video based on that instruction. This is suitable for creating a concept clip, a simple social post, or an initial product-video draft without supplying source media.

First-and-last-frame generation

The model can use up to one first-frame image and one last-frame image to guide how a video begins and ends. This is useful when the transition itself matters—for example, showing a product changing position, moving from one scene to another, or transforming between two visual states.

The two images act as visual anchors rather than merely decorative references. The generated frames between them are still produced by the model, so the feature should be treated as guided transition generation rather than a guarantee of exact object or motion continuity.

Reference-based video generation

Reference-based generation allows a prompt to be combined with reference images, videos, and audio. These inputs can provide guidance about appearance, movement, characters, scenes, or sound. Tencent Cloud recommends using at least one reference image or video for this mode and keeping the total reference assets within the documented limits.

This makes the Lite model more suitable than a text-only workflow when a campaign needs to preserve a product appearance, draw on an existing motion sample, or coordinate several kinds of creative direction. The supplied research does not specify a fixed maximum number of assets, so implementations should follow the limits in the current Tencent Cloud API documentation rather than assume an undocumented allowance.

Supported specifications

SpecificationDocumented support
Model IDwand-vega-video-1.0-lite
Video duration4 to 15 seconds; 5 seconds by default
Resolution768P, 1080P, 2K, and 4K
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and adaptive
InputsText, image URLs, video URLs, and audio URLs
OutputGenerated video URL, with an optional last-frame image
ProcessingAsynchronous task submission and polling
Default concurrencyFive, according to the listed model information

The available resolution range is unusually broad for a lightweight short-video model, extending from 768P through 4K. However, supported values are discrete: a request outside the documented range is rejected rather than automatically downgraded. Teams should validate resolution and aspect-ratio choices before submitting production jobs.

How the TokenHub workflow operates

WAND-Vega-Video-1.0-Lite is accessed through Tencent Cloud TokenHub using a task-based video endpoint. An application submits the model ID, a content array containing the prompt and any reference media, and optional generation settings such as duration, resolution, aspect ratio, and output options.

The service returns a task identifier rather than an immediately available video. The application then polls the task until it succeeds or fails. This asynchronous design fits batch production, where a queue can manage many jobs, but it is less convenient for an interface that expects a video to appear immediately.

Successful video URLs are temporary and are documented as remaining valid for 12 hours. Production systems should download each result and move it to durable storage promptly. A robust integration should also handle queued tasks, failures, retries, output transfer, and expiration of unused links.

Strengths and trade-offs

The model’s clearest strength is the combination of short-video generation and multiple reference modalities. A team can start with a text-only concept, add opening and closing frames for a controlled transition, or use images, video, and audio to provide richer direction. The range of durations, aspect ratios, and resolutions also covers common advertising, e-commerce, and social-media formats.

Its lightweight positioning is another practical advantage for high-volume workflows. The supplied editorial assessment rates its speed and cost characteristics at 8 out of 10. Those scores are editorial evaluations, not Tencent Cloud benchmarks or provider-published guarantees. They indicate that the model appears well suited to fast, cost-conscious production, while the exact cost and latency for a workload still depend on service conditions and request settings.

The trade-off is scope. The model generates short clips rather than long-form video, and it does not provide a documented streaming-generation interface. Its asynchronous operation adds queue management and polling work. Higher resolution, longer duration, or more complex reference inputs may also affect practical throughput, but the supplied documentation does not provide verified latency or per-generation performance figures.

Pricing and limits

Tencent Cloud’s current public pricing information does not expose a verified model-specific price for WAND-Vega-Video-1.0-Lite. There is therefore no reliable price-per-video figure to report. Prospective users should confirm the applicable TokenHub billing terms directly before estimating campaign costs.

The published material also does not specify a knowledge cutoff, text context window, maximum text-token output, fine-tuning support, batch API, or a separate release date for this exact model. These omissions are important when evaluating the model: they mean those capabilities should be treated as unknown, not as supported by default.

Some limits are explicit. The generated duration must be between 4 and 15 seconds, supported resolution values must be selected, and a text element is required. The video output link is temporary. Other operational limits, including the precise maximum number of reference assets, should be checked against the current API guide.

Modalities and capability profile

WAND-Vega-Video-1.0-Lite accepts text, image, video, and audio inputs and produces video output. It can optionally return a last-frame image alongside the generated video. In practical terms, it is a multimodal video model: the media inputs influence a visual or audiovisual generation task rather than being used for open-ended conversation.

CapabilityAssessment
Text inputSupported and required
Image inputSupported through image references and frame controls
Video inputSupported as a reference input
Audio inputSupported as a reference input
Video outputSupported
Text outputNot the model’s documented output type
Reasoning or codingNot documented as model capabilities
Tool or function callingNot documented
Streaming outputNot supported; generation is asynchronous

These distinctions matter when selecting the model for an application. WAND-Vega-Video-1.0-Lite should be paired with separate software or models for prompt construction, workflow control, metadata generation, content moderation, or other text-based tasks unless those functions are implemented by the surrounding application.

Best use cases

  • E-commerce product clips: Generate multiple short demonstrations or promotional variations from product imagery and a text brief.
  • Social and advertising assets: Produce vertical, square, widescreen, or other aspect-ratio variants for short-form campaigns.
  • Image-to-video animation: Turn a still product, character, or scene into a moving clip.
  • First-to-last-frame transitions: Create a controlled visual change between two supplied states.
  • Reference-guided production: Combine images, sample video, and audio with a prompt when a text-only instruction is not specific enough.
  • Batch generation: Submit many asynchronous tasks for later review and storage.

When to choose this model

Choose WAND-Vega-Video-1.0-Lite when the requirement is a short generated clip, especially when production volume, multiple reference types, and cost awareness matter more than real-time response or long-duration continuity. It is a reasonable fit for teams building a queued content pipeline rather than a live interactive video-generation experience.

Another option may be more appropriate when the project needs long-form video, documented streaming generation, fine-tuning, a batch API with formally specified behavior, or a published model-specific price that can be used directly for budgeting. A different model may also be preferable when the workflow requires text reasoning, coding, tool calls, or structured text output in the same model. Those requirements are not documented for WAND-Vega-Video-1.0-Lite.

For evaluation, begin with representative prompts and reference media from the intended campaign. Test the actual duration, resolution, aspect-ratio, polling, retry, and storage workflow together; a model that produces suitable clips is only operationally useful if the surrounding system can preserve its temporary outputs and handle asynchronous completion reliably.


Answers to Frequently Asked Questions

Is pricing available for WAND-Vega-Video-1.0-Lite?
A verified model-specific public price is not provided in the available Tencent Cloud information. Users should confirm the current TokenHub billing terms directly before estimating production costs. Exact latency and operational limits, including the maximum number of reference assets, should also be checked in the current API documentation.
How does the WAND-Vega-Video-1.0-Lite TokenHub API work?
The application submits an asynchronous task through Tencent Cloud TokenHub using the model ID, a prompt, optional reference media, and generation settings. The service returns a task identifier, which the application polls until completion or failure. Successful video URLs are temporary and remain valid for 12 hours, so results should be downloaded to durable storage promptly.
What generation modes does WAND-Vega-Video-1.0-Lite support?
The model supports text-to-video, first-and-last-frame generation, and reference-based video generation. Every request must include at least one text element, while reference-based workflows can also use image, video, and audio inputs.
What are the supported duration, resolution, and aspect-ratio options?
Generated videos can be 4 to 15 seconds long, with 5 seconds as the default. Supported resolutions are 768P, 1080P, 2K, and 4K. Supported aspect ratios are 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and adaptive.
What is WAND-Vega-Video-1.0-Lite used for?
WAND-Vega-Video-1.0-Lite is a lightweight Tencent Cloud video-generation model for creating short clips from text prompts and optional reference images, videos, or audio. It is suited to product demonstrations, advertising variations, social-media content, image-to-video animation, and batch content production.


Sources 3
Provider

About Tencent AI