Veo 3.1

Veo 3.1 Lite

by Google DeepMind · Preview; currently available through the Gemini API and Google Cloud Vertex AI

Veo 3.1 Lite is Google DeepMind's cost-focused preview model for short text-to-video and image-to-video generation. It produces 4-, 6-, or 8-second 720p or 1080p clips with native audio, supports 16:9 and 9:16 formats, and costs $0.05 per second at 720p or $0.08 per second at 1080p through the Gemini API.

Video generation Audio Reasoning Coding
Veo 3.1 Lite is a developer-focused video model from Google DeepMind. It turns text prompts or still images into short videos with synchronized audio, offering a lower per-second price than the other Veo 3.1 tiers. Its main appeal is economical, repeatable generation for social content, advertising variations, storyboards, product demonstrations, and other workflows where many clips are more useful than maximum production capability.
Outputs

What Veo 3.1 Lite can produce

Video generation Audio
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Veo 3.1
Model type Other
Release date 2026-04-03
Status Preview; currently available through the Gemini API and Google Cloud Vertex AI
Knowledge cutoff notes

Google does not publish a conventional textual knowledge-cutoff date for Veo 3.1 Lite. It is a generative video model rather than a general-purpose knowledge model.

Model notes

Google DeepMind describes Veo 3.1 Lite as an addition to the Veo 3 series. The Gemini API preview identifier is veo-3.1-lite-generate-preview. It accepts text and images and returns one video with audio per request. Supported output resolutions are 720p and 1080p; 4K is not supported. Supported durations are 4, 6, and 8 seconds, with 8 seconds required for 1080p. The Gemini API documentation says that video extension is unavailable for Veo 3.1 Lite. Google model-card evaluations reported a 54.6% overall win rate for text-to-video across 1,000 prompts and a 47.2% overall win rate for image-to-video across 646 prompts against Veo 3.1 Fast. Editorial reasoning and coding scores are not meaningful for this specialist video model and are set to the minimum comparative value.

Cost

Model pricing

Input $0.05 per second for 720p video with audio; $0.08 per second for 1080p video with audio
Output Video with native audio, priced per generated-video second
Model guide

Veo 3.1 Lite: Cost-Efficient Video Generation with Native Audio

Veo 3.1 Lite is Google DeepMind's preview video-generation model for high-volume text-to-video and image-to-video workflows. It creates four-, six-, or eight-second 720p or 1080p videos with native audio, supports 16:9 and 9:16 output, and is priced at $0.05 per second for 720p or $0.08 per second for 1080p through the Gemini API.

What is Veo 3.1 Lite?

Veo 3.1 Lite is a preview video-generation model provided by Google DeepMind. It belongs to the Veo 3.1 family and is designed for developers, enterprises, and creative teams that need to create short video clips at scale while keeping generation costs under control.

The model accepts a written prompt or an input image. A prompt can describe the subject, setting, movement, camera direction, dialogue, sound effects, and ambient audio. An image can serve as the visual starting point for an animated clip. The result is a video file with natively generated audio rather than a silent sequence that requires a separate audio-generation step.

Veo 3.1 Lite is not a general-purpose conversational or coding model. It is a specialist generative-video system, so its value should be judged by the quality, cost, and practical limits of its video output rather than by language-model benchmarks.

Where it fits in the Veo lineup

Google positions Veo 3.1 Lite as the most cost-effective member of the Veo 3.1 family. The supplied documentation describes it as an addition to the Veo 3 series and makes it available as a preview through the Gemini API and Google Cloud Vertex AI. It is also distributed through product channels including Gemini AI Studio and Google Flow.

Its positioning is straightforward: Veo 3.1 Lite favors generation economics and volume over the broadest feature set. The standard Veo 3.1 model is the more appropriate direction when a project needs higher-end production capabilities, advanced reference-image control, video extension, or the strongest available final-production quality. Veo 3.1 Fast may be a better choice when reducing latency matters more than achieving the lowest possible per-second price.

These are positioning trade-offs rather than a claim that Lite is universally better or worse. The right choice depends on whether the workflow is constrained primarily by cost, waiting time, creative controls, or output quality.

Supported inputs and output

Veo 3.1 Lite supports two documented generation workflows:

  • Text-to-video: a natural-language description is converted into a video.
  • Image-to-video: an input image is animated or used as the starting visual composition.

Its documented output characteristics are:

SpecificationSupported value
Output typeVideo with native audio
Resolution720p or 1080p
Aspect ratios16:9 and 9:16
Duration4, 6, or 8 seconds
Frame rate24 frames per second
Videos per requestOne

Eight-second output is required when generating 1080p video. The model does not support 4K output according to the supplied documentation. The documented Gemini API workflow also does not support video extension, video-to-video input, or reference-image workflows for Veo 3.1 Lite.

Native audio capabilities

A key distinction is that Veo 3.1 Lite generates audio together with the video. Depending on the prompt, that audio can include dialogue, sound effects, and ambient sound. For example, a prompt might request a short product demonstration with a spoken line, room ambience, and the sound of an object being placed on a table.

Native audio can simplify short-form production because the visual and audio tracks are created in one generation request. It does not make the model a standalone speech, music, or audio-production system, however. Its primary output remains a short generated video, and the supplied research does not identify dedicated music-generation or standalone speech-output capabilities.

API access and model identifier

The Gemini API preview identifier is veo-3.1-lite-generate-preview. The model is also available through Google Cloud Vertex AI, subject to the relevant Google Cloud access, availability, and pricing arrangements.

Because the model is in preview, its behavior, availability, pricing, and supported features may change. Developers should use the current Google documentation and the exact preview identifier rather than assuming that a standard Veo 3.1 or Veo 3.1 Fast endpoint is interchangeable with Lite.

The model is suited to asynchronous generation workflows in which an application submits a prompt, waits for the generated media, and then stores or processes the resulting video. The supplied research does not document streaming output, function calling, tool use, structured JSON output, or batch API support for this model.

Pricing and cost examples

Gemini API pricing for Veo 3.1 Lite is calculated per second of generated video. The listed rates are:

  • 720p with audio: $0.05 per generated second
  • 1080p with audio: $0.08 per generated second

Using those rates, an eight-second clip costs approximately $0.40 at 720p or $0.64 at 1080p, before any additional platform-specific charges. A four-second 720p clip would cost approximately $0.20, while a six-second 1080p clip would cost approximately $0.48. These are arithmetic examples based on the published per-second rates, not separate plans or guaranteed final invoices.

No free tier is listed for Veo 3.1 Lite in the current Gemini API pricing information supplied for this page. Developers should also account for retries, failed or unwanted generations, storage, and any charges associated with the platform through which the model is accessed.

Strengths and trade-offs

Strengths

  • Lower generation cost: its per-second price is designed for high-volume use and is lower than the other Veo 3.1 tiers described in the supplied research.
  • Native audio: dialogue, effects, and ambient sound can be generated with the video.
  • Two useful input modes: teams can start from either a text concept or a still image.
  • Portrait and landscape output: 16:9 and 9:16 formats support conventional video and vertical social content.
  • 1080p availability: the model can produce higher-resolution output when an eight-second clip is acceptable.
  • Short-clip iteration: limited durations can be useful for testing concepts, generating variations, and assembling larger edits from multiple clips.

Limitations

  • Preview status: specifications and availability may change.
  • Short maximum duration: the documented choices stop at eight seconds.
  • No 4K: it is not intended for workflows that require 4K generation.
  • No documented video extension: users cannot rely on the Gemini API Lite workflow to continue an existing generated clip.
  • No documented video-to-video workflow: the supported input modes are text and images, not an existing video.
  • No documented reference-image workflow: the Lite documentation does not list the more advanced reference-image control available in some higher-tier workflows.
  • One output per request: applications that need multiple alternatives must submit multiple generations and pay for the resulting video seconds.

Reasoning, coding, and tool support

Veo 3.1 Lite is not designed to perform general reasoning, software development, or conversational assistance. The supplied comparative data assigns minimum editorial scores for reasoning and coding because those categories are not meaningful measures for a specialist video model; those scores are editorial classifications, not provider-published benchmark results.

No general-purpose tool or function-calling capability is documented for the model. It can interpret instructions contained in a video prompt, but that should not be confused with using external tools, browsing the web, executing code, or returning structured application data. The documented output is generated video with audio.

When to choose Veo 3.1 Lite

Choose Veo 3.1 Lite when the workflow needs many short clips and the cost per generated second is a major consideration. Suitable examples include:

  • Social-media video variations in both landscape and portrait formats
  • Advertising concept tests with multiple visual treatments
  • Storyboards and previsualization for film or commercial projects
  • Short product demonstrations and promotional clips
  • Educational or instructional segments
  • Automated media pipelines that generate large numbers of brief assets
  • Rapid experimentation with image-to-video animation

It is particularly practical when an eight-second 720p clip is sufficient, because that combination provides the lowest listed cost for the longest documented duration. Use 1080p when the additional resolution justifies the higher per-second price and the workflow can accept the eight-second requirement.

When another option may be better

Veo 3.1 is the more suitable option when the project depends on capabilities that are not documented for Lite, such as higher-end production quality, advanced reference-image control, or video extension. Veo 3.1 Fast may be preferable when users need lower latency and are willing to trade away the lowest possible generation cost.

A conventional editing or compositing workflow may also be more appropriate when the project requires long-form continuity, precise shot-by-shot control, reliable character consistency, or extensive post-production. Veo 3.1 Lite creates short generated clips; it is not presented as a complete replacement for an editing suite or a system for generating long continuous scenes.

Bottom line

Veo 3.1 Lite is a cost-focused preview model for short video generation with native audio. Its combination of text-to-video and image-to-video input, 720p or 1080p output, portrait and landscape formats, and per-second pricing makes it well suited to scalable creative experimentation. Its limits are equally important: clips are short, only one video is returned per request, 4K and several advanced video workflows are unavailable, and preview behavior may change. It is best understood as an economical production component for high-volume short-form video rather than as the most capable or fastest option in the Veo family.


Answers to Frequently Asked Questions

What is the Veo 3.1 Lite API model identifier?
The Gemini API preview identifier for Veo 3.1 Lite is "veo-3.1-lite-generate-preview". The model is also available through Google Cloud Vertex AI, subject to applicable access, availability, and pricing arrangements.
Does Veo 3.1 Lite generate audio?
Yes. Veo 3.1 Lite generates native audio with the video, including dialogue, sound effects, and ambient sound when requested in the prompt. It is primarily a video-generation model rather than a standalone speech, music, or audio-production system.
How much does Veo 3.1 Lite cost?
Gemini API pricing is $0.05 per generated second for 720p video with audio and $0.08 per generated second for 1080p video with audio. For example, an eight-second clip costs approximately $0.40 at 720p or $0.64 at 1080p, before any additional platform charges.
What video formats and durations does Veo 3.1 Lite support?
Veo 3.1 Lite supports 720p and 1080p resolution, 16:9 and 9:16 aspect ratios, and clip durations of 4, 6, or 8 seconds at 24 frames per second. Eight-second output is required for 1080p generation, and the model does not support 4K output.
What is Veo 3.1 Lite?
Veo 3.1 Lite is a Google DeepMind preview model for generating short videos from text prompts or input images. It is designed for cost-efficient, high-volume video production and generates native audio together with the video.


Sources 5
Provider

About Google DeepMind