Veo 3.1

Veo 3.1

by Google DeepMind · Preview; currently accessible through Google APIs and Google products

Veo 3.1 is Google DeepMind’s specialized video-generation model for short text-to-video and image-to-video clips with synchronized native audio. It supports reference images, video extension, first-and-last-frame transitions, 16:9 and 9:16 formats, 4-, 6-, and 8-second outputs, 720p and 1080p resolution, and supported 4K workflows. Gemini API pricing is listed per generated second, while spoken dialogue and long-form production remain limitations.

Video generation Audio Reasoning Coding
Veo 3.1 is Google DeepMind’s current Veo video-generation model line, designed for cinematic text-to-video and image-to-video creation with synchronized generated audio. It is aimed at short-form storytelling, advertising concepts, storyboarding, visual-effects exploration, and creative previsualization rather than conversational language or general-purpose coding.
Outputs

What Veo 3.1 can produce

Video generation Audio
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
5/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family Veo 3.1
Model type Multimodal
Context window 1K tokens
Release date 2025-10-15
Status Preview; currently accessible through Google APIs and Google products
Model notes

The canonical model family is Veo 3.1. Gemini API identifiers include veo-3.1-generate-preview, veo-3.1-fast-generate-preview, and veo-3.1-lite-generate-preview. The standard Veo 3.1 model supports text-to-video, image-to-video, video-to-video or video extension workflows, reference images, first-and-last-frame generation, 16:9 and 9:16 aspect ratios, 24 fps, and 4-, 6-, or 8-second clips. Supported output resolutions include 720p and 1080p, with 4K available in supported 8-second workflows. Native audio generation is always enabled. The Gemini API documents Veo 3.1 as preview and lists a 1,024-token text input limit with one video output per request. Google Cloud documentation also lists veo-3.1-generate-001 as a stable Vertex AI endpoint corresponding to the preview generation model. Google states that spoken audio, especially short dialogue segments, remains an area of active development. Editorial scores are comparative estimates for a video-generation model and are not vendor-published ratings.

Cost

Model pricing

Input $0.40 per second for standard Veo 3.1 video with audio at 720p or 1080p; $0.60 per second at 4K
Output Generated video with native audio; pricing is charged per generated video second
Model guide

Veo 3.1: Google DeepMind’s Video Model with Native Audio and Reference Controls

Veo 3.1 is Google DeepMind’s video-generation model for creating short clips from text, images, and video inputs. Its defining capabilities include synchronized native audio, improved prompt adherence, reference-image controls, scene extension, first-and-last-frame transitions, vertical video, and supported 1080p and 4K workflows.

What is Veo 3.1?

Veo 3.1 is Google DeepMind’s video-generation model for producing short clips from written prompts and visual references. Unlike a text-only model, it generates video as its primary output and can also create synchronized audio, including sound effects, ambience, music, and attempted spoken dialogue. Google describes it as part of the current Veo model line and positions it for cinematic creation across Google products and developer platforms.

The model can turn a description such as “a slow tracking shot through a rain-soaked neon market at night” into a short video. It can also use an input image as a visual starting point, accept reference images to guide elements such as characters or objects, extend an existing video, or generate a transition between specified first and last frames. These controls make it more useful for planned visual sequences than a basic one-prompt video generator.

Veo 3.1 was released on October 15, 2025, according to the supplied launch research. Its Gemini API identifiers include veo-3.1-generate-preview, veo-3.1-fast-generate-preview, and veo-3.1-lite-generate-preview. The Gemini API documentation describes the model family as preview, while Google Cloud documentation also lists veo-3.1-generate-001 as a stable Vertex AI endpoint corresponding to the generation model. Availability and exact limits can therefore depend on the Google product or endpoint being used.

Where Veo 3.1 fits in Google’s lineup

Veo 3.1 sits within Google DeepMind’s generative video offering and is exposed through several parts of Google’s ecosystem, including the Gemini API, Vertex AI, Google AI Studio, Flow, Gemini, Google Vids, and related Google products. This gives it two broad audiences: creators using a Google application and developers or production teams integrating video generation into an application or workflow.

The model is more specialized than a general multimodal assistant. Its main job is not to answer questions, write software, or analyze documents. Instead, it converts visual direction into short generated video. The surrounding Google products may provide the interface, editing workflow, or account and billing system, but Veo 3.1 remains the video-generation component.

Core video and audio capabilities

Veo 3.1 supports both text-to-video and image-to-video generation. A text prompt can specify the subject, setting, camera movement, lighting, style, action, and timing. An image prompt can establish an initial composition or visual identity before the model animates it. The supplied specifications also list video input support, which is relevant to video-to-video or video-extension workflows.

  • Text-to-video: creates a clip from a written description.
  • Image-to-video: animates or develops an uploaded image.
  • Video extension: continues an existing sequence in supported workflows.
  • Reference images: helps guide recurring visual elements, characters, or objects.
  • First-and-last-frame generation: creates a transition constrained by specified beginning and ending frames.
  • Aspect ratios: supports 16:9 landscape and 9:16 vertical output.
  • Frame rate: supports 24 frames per second.
  • Audio: native audio generation is enabled, rather than requiring a separate audio-generation step.

Native audio is one of Veo 3.1’s most important practical distinctions. A generated clip can include audio synchronized with the visual action, which is useful for previews, social-video concepts, advertisements, and storyboards that need a rough sound layer. However, Google states that spoken audio, particularly short dialogue segments, remains an area of active development. Users should treat generated speech as an experimental component rather than assuming consistently accurate pronunciation, wording, or timing.

Output durations, resolution, and limits

The documented clip durations are 4, 6, or 8 seconds. This makes Veo 3.1 suitable for generating shots and short sequences, but not for producing a complete long-form film in one request. Longer projects need to be assembled from multiple generations, with the usual challenges of maintaining character identity, continuity, camera logic, and audio consistency between shots.

Supported output resolutions include 720p and 1080p. The research also identifies 4K output in supported 8-second workflows, so 4K should not be treated as universally available for every duration, interface, or endpoint. Vertical 9:16 output is useful for mobile-first content, while 16:9 is better suited to conventional landscape video and presentation formats.

The Gemini API lists a 1,024-token text input limit and one video output per request. The token limit applies to the textual instruction rather than representing a maximum film script or a general conversational context window. A prompt can be detailed, but a very long screenplay or large set of instructions may need to be condensed into the most important visual and audio directions.

Supported input and output types

Veo 3.1 is multimodal in the specific sense that it can work with text, images, and video inputs. Its primary output is video, with generated audio accompanying the video in supported generation workflows. It is not a text-output model, image-output model, or general-purpose speech model in the supplied specification.

CapabilityVeo 3.1 support
Text inputYes, with a documented 1,024-token text input limit
Image inputYes
Video inputYes, for supported extension or video-to-video workflows
Audio inputNot listed as supported
Video outputYes
Audio outputYes, generated natively with the video
Text or structured JSON outputNot the model’s output purpose

There is no documented tool or function-calling capability for Veo 3.1 itself. An application can combine the model with other services, storage, editing tools, or workflow automation, but those integrations are outside the model’s native generation capability.

Veo 3.1 pricing

The supplied Gemini API pricing research lists standard Veo 3.1 video with audio at $0.40 per generated second at 720p or 1080p. The listed price for 4K output is $0.60 per generated second. Charges are based on the generated video duration, not on the number of words in the prompt.

Using those rates, an 8-second 720p or 1080p generation would cost approximately $3.20, while an 8-second 4K generation would cost approximately $4.80, before any applicable platform-specific terms or taxes. These examples describe the arithmetic of the supplied per-second rates, not a promise that every Google product uses identical billing or exposes every resolution.

Google also documents faster and lighter Veo 3.1 identifiers, including veo-3.1-fast-generate-preview and veo-3.1-lite-generate-preview. The supplied research does not provide separate prices or a complete capability comparison for those variants, so their cost and quality trade-offs should be checked in the relevant Google documentation rather than assumed.

Main strengths and limitations

Strengths

  • Integrated audio generation: video and sound can be created in the same generation workflow.
  • Useful visual control: reference images, first-and-last-frame transitions, and scene extension provide more direction than a basic text-only prompt.
  • Flexible composition: 16:9 and 9:16 formats support both landscape and vertical publishing contexts.
  • Multiple access points: creators and developers can encounter the model through Google applications, Google AI Studio, the Gemini API, or Vertex AI.
  • High-resolution options: 1080p is supported, with 4K available in specified 8-second workflows.

Limitations

  • Short clips: individual outputs are limited to 4, 6, or 8 seconds, so longer work requires multiple shots.
  • Audio uncertainty: spoken dialogue remains an active development area and may not be reliable enough for final production.
  • Preview status: Gemini API access is documented as preview, which means behavior, limits, availability, or identifiers may change.
  • Prompt-size constraint: the documented 1,024-token text input limit restricts how much textual direction can be sent in one request.
  • Non-deterministic rendering: generated visuals and audio should not be expected to reproduce a perfectly controlled result on every attempt.
  • Specialized purpose: it is not an appropriate choice for ordinary chat, software development, structured data generation, or long-form text reasoning.

Reasoning, coding, and speed trade-offs

Veo 3.1 does not provide conventional language-model reasoning or coding as its primary function. The supplied editorial data assigns both reasoning and coding scores of 1, but these are comparative editorial estimates for a video-generation model, not ratings published by Google. In practical terms, the model should be evaluated by the quality and controllability of its generated video rather than by asking it to solve programming problems or conduct extended analysis.

The model is also comparatively expensive and slower to use than a text-generation model because each request produces a rendered audiovisual artifact. The supplied editorial speed score is 5 and cost score is 4; these are likewise subjective comparative estimates, not provider benchmarks. The practical cost trade-off is clearest when choosing between a standard generation and a 4K generation: 4K costs more per second, while 720p or 1080p may be more economical for concept exploration and iteration.

Best use cases

Veo 3.1 is a strong fit when a project needs short, visually directed shots and an initial synchronized sound layer. Suitable examples include:

  • storyboards and cinematic previsualization;
  • advertising and campaign concepts;
  • short social-video sequences in vertical format;
  • visual-effects exploration before expensive production work;
  • product or environment concept videos;
  • short-form storytelling assembled from multiple generated shots;
  • testing camera movement, lighting, composition, or scene transitions.

Reference images are especially useful when the creator wants to maintain a particular look across a generation or guide the model toward a specific object, person, or setting. First-and-last-frame generation can help when the important requirement is a controlled visual transition rather than an unconstrained clip.

When to choose Veo 3.1

Choose Veo 3.1 when native audio, cinematic short clips, image guidance, or video-extension controls matter more than conversational reasoning. It is particularly appropriate for creators who already work within Google’s ecosystem or developers who want to access video generation through the Gemini API or Vertex AI.

Another option may be more appropriate when the project needs long continuous footage, deterministic frame-by-frame control, dependable spoken dialogue, or detailed post-production control over every sound and edit. A general-purpose language model is a better choice for scripting, planning, coding, or prompt development. A dedicated editing or production pipeline is better for assembling many shots into a finished long-form video. Veo 3.1’s value is in generating short audiovisual building blocks, not replacing the entire production process.

Bottom line

Veo 3.1 is a specialized Google DeepMind video model distinguished by native audio generation and a growing set of visual-control features. Text, images, and supported video workflows can be used to create 4-, 6-, or 8-second clips in landscape or vertical formats, with 720p and 1080p output and 4K available in supported 8-second workflows. Its main trade-offs are short output duration, preview-era availability, per-second cost, and still-developing spoken audio. For cinematic ideation, short-form content, reference-guided animation, and audiovisual previsualization, those trade-offs may be worthwhile; for general reasoning or long-form production, they make another tool a better fit.


Answers to Frequently Asked Questions

How much does Veo 3.1 cost?
The supplied Gemini API pricing lists standard Veo 3.1 video with audio at $0.40 per generated second in 720p or 1080p, and $0.60 per generated second in 4K. At these rates, an 8-second clip costs approximately $3.20 at 720p or 1080p and $4.80 at 4K.
What are the main limitations of Veo 3.1?
Veo 3.1 is limited to short 4-, 6-, or 8-second clips, has a documented 1,024-token text input limit, and may produce inconsistent spoken dialogue. Gemini API access is documented as preview, availability and limits may vary by Google product or endpoint, and generated visuals and audio are not perfectly deterministic.
How long and how high-resolution are Veo 3.1 videos?
Veo 3.1 generates clips lasting 4, 6, or 8 seconds. Supported resolutions include 720p and 1080p, while 4K is available in supported 8-second workflows. Longer videos must be assembled from multiple generated clips.
What is Veo 3.1?
Veo 3.1 is Google DeepMind’s video-generation model for creating short clips from text prompts, images, and supported video inputs. It can generate synchronized audio, including sound effects, ambience, music, and experimental spoken dialogue.
What input and output formats does Veo 3.1 support?
Veo 3.1 supports text, image, and video inputs for workflows such as text-to-video, image-to-video, video extension, reference-guided generation, and first-and-last-frame transitions. It produces video with natively generated audio and supports 16:9 and 9:16 aspect ratios at 24 frames per second.


Sources 6
Provider

About Google DeepMind