Veo 3.1

Veo 3.1 Fast

by Google DeepMind · Generally available on Vertex AI; preview model on the Gemini API

Veo 3.1 Fast is Google DeepMind's speed-optimized video model for short audiovisual clips. It supports text and image inputs, native audio, 16:9 and 9:16 formats, first-and-last-frame guidance, four-to-eight-second durations, and per-second Gemini API pricing. It is best for rapid creative iteration and high-volume production rather than long-form video or maximum-fidelity final renders.

Video generation Audio
Veo 3.1 Fast is a Google DeepMind video-generation model designed for applications where rapid iteration and per-second cost matter more than the highest available Veo quality. It turns text prompts and supported image inputs into short 24-frame-per-second videos with automatically generated dialogue, sound effects, and ambient audio. The model is available as a preview through the Gemini API and as a production model on Vertex AI, although supported features and model identifiers differ between those environments.
Outputs

What Veo 3.1 Fast can produce

Video generation Audio
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Veo 3.1
Model type Video Generation
Context window 1K tokens
Release date 2025-10-15
Status Generally available on Vertex AI; preview model on the Gemini API
Knowledge cutoff notes

Google does not publish a conventional knowledge-cutoff date for Veo 3.1 Fast. The model generates video from supplied prompts and media inputs rather than operating as a general-purpose knowledge model.

Model notes

The Gemini API model identifier is veo-3.1-fast-generate-preview. Vertex AI documents veo-3.1-fast-generate-001 as the production model identifier. The Gemini API documentation describes the model as preview, while Google Cloud announced Veo 3.1 and Veo 3.1 Fast as generally available on Vertex AI in November 2025. It accepts text and image inputs and generates short videos with audio. Supported durations are 4, 6, or 8 seconds, with 24 FPS and 16:9 or 9:16 aspect ratios. Gemini API documentation lists 720p, 1080p, and a 4K pricing tier; the Vertex AI fast endpoint documentation lists 720p and 1080p. Reference-image and video-extension support varies by endpoint. The context value represents the documented 1,024-token text prompt limit, not a conventional language-model context window. Reasoning and coding scores are not applicable to this video-generation model.

Cost

Model pricing

Input $0.10 per video second at 720p; $0.12 per video second at 1080p; $0.30 per video second at 4K where supported
Output Video with natively generated audio, priced per generated video second
Model guide

Veo 3.1 Fast: Faster, Lower-Cost Video Generation with Native Audio

Veo 3.1 Fast is Google DeepMind's speed-optimized video generation model for creating short clips from text and images. It generates video with synchronized native audio, supports portrait and landscape formats, first-and-last-frame guidance, 720p and 1080p output, and four-, six-, or eight-second durations. Its main advantage over standard Veo 3.1 is faster, less expensive iteration, while its trade-offs include short output duration, endpoint-specific feature differences, preview status in the Gemini API, and lower positioning for maximum-fidelity final production.

What is Veo 3.1 Fast?

Veo 3.1 Fast is the speed-optimized member of Google DeepMind's Veo 3.1 video-generation family. It produces short audiovisual clips rather than text responses: a request can result in a video with synchronized visual content and natively generated audio. That audio may include spoken dialogue, sound effects, and environmental ambience when those elements are described in the prompt.

The model is aimed at high-volume creation and fast experimentation. Compared with using a higher-quality video model for every draft, Veo 3.1 Fast is positioned for workflows such as testing advertising concepts, generating social-media variations, exploring storyboards, and producing automated video content. It remains a short-form model, so it is not designed to create long films or extended continuous scenes in a single request.

Google DeepMind announced the model for public preview in the Gemini API on October 15, 2025. Google Cloud documentation identifies veo-3.1-fast-generate-001 as the Vertex AI production model, while the Gemini API uses veo-3.1-fast-generate-preview. The Gemini API documentation still describes the latter as a preview model.

Where it fits in the Veo lineup

Veo 3.1 Fast is not a general-purpose language model or a lower-cost chat model. It is a specialized video generator within the Veo 3.1 family. Its defining trade-off is speed and generation cost: it is intended to make more iterations possible when the user does not need the highest-fidelity configuration for every render.

The standard Veo 3.1 model is the more appropriate comparison when visual fidelity is the primary concern. Veo 3.1 Fast is a better fit when a team needs to explore many concepts, create multiple short variations, or keep automated generation costs under control. Google also describes Veo 3.1 Lite as a more cost-sensitive alternative in scenarios where its narrower supported feature set is sufficient. These comparisons are positioning guidance rather than a published benchmark ranking; the supplied research does not provide numerical quality or latency benchmarks between the models.

Inputs and generation controls

Veo 3.1 Fast accepts text prompts and image inputs. Text-to-video generation creates a clip from a written description, while image-to-video generation animates or develops supplied visual material. Depending on the endpoint, developers can also provide first-frame and last-frame images to guide how a clip begins and ends.

  • Text input: prompts describing the scene, movement, style, dialogue, sound, or other desired details.
  • Image input: source imagery for image-to-video generation.
  • Frame guidance: first-frame and last-frame images where supported by the selected endpoint.
  • Aspect ratios: landscape 16:9 and portrait 9:16.
  • Durations: four, six, or eight seconds.
  • Frame rate: 24 frames per second.

Gemini API documentation also describes video-to-video support for the broader Veo 3.1 and Veo 3.1 Fast family. However, endpoint documentation is not identical: the Vertex AI fast production endpoint lists a narrower set of operations and does not list reference-image-to-video or Veo-video extension support. Developers should therefore check the exact API, model identifier, and region instead of assuming that every Veo 3.1 Fast capability is available everywhere.

Video, audio, and output limits

The model generates one video output per request in the Gemini API. Supported output resolutions documented for Veo 3.1 Fast include 720p and 1080p. The Gemini API pricing page also lists a 4K price tier, but the model-specific capability documentation identifies 4K most clearly for the broader Veo 3.1 family. Availability of 4K for Veo 3.1 Fast can therefore depend on the endpoint, and production systems should verify it before building around that resolution.

Audio is always enabled for Veo 3.1 Fast. A prompt can request spoken lines, sound effects, or ambient audio alongside the visual action. This makes the model different from video generators that return silent footage and require a separate audio workflow. Audio-processing checks can nevertheless prevent a requested generation, and the model's output is watermarked with SynthID.

The documented text prompt limit is 1,024 tokens. This is a prompt-length limit rather than a conventional language-model context window. The research does not specify a separate maximum number of output tokens because the model's primary output is video. Its practical output limit is expressed through clip duration, resolution, and the one-video-per-request behavior documented for the Gemini API.

Pricing and access

In the Gemini API, Veo 3.1 Fast is listed as a paid preview model priced by generated video second. The documented rates are:

Output tierPrice per generated video second
720p$0.10
1080p$0.12
4K where supported$0.30

At those rates, a four-second 720p clip would cost $0.40 and an eight-second 1080p clip would cost $0.96, before any account-specific billing considerations. These examples apply the published per-second prices and are not a separate subscription price. There is no free tier for this model in the supplied Gemini API research.

Access and feature availability differ between the Gemini API and Vertex AI. The Gemini API uses the preview identifier veo-3.1-fast-generate-preview, while Vertex AI uses veo-3.1-fast-generate-001 for the production endpoint. Google Cloud announced Veo 3.1 and Veo 3.1 Fast as generally available on Vertex AI in November 2025. This means that “preview” and “generally available” describe different service environments rather than a single universal status.

Main strengths and trade-offs

The clearest strength of Veo 3.1 Fast is its suitability for repeated generation. A lower per-second price than standard Veo 3.1, combined with a speed-focused configuration, makes it practical to test several prompts or produce many short variations. Native audio is another important advantage when a workflow needs a complete audiovisual clip rather than silent video that must be edited later.

  • Fast iteration: useful for drafts, concept exploration, and high-volume generation.
  • Lower generation cost: the published Gemini API rates are below the listed standard-quality positioning for workflows that prioritize economy.
  • Native audio: dialogue, sound effects, and ambience can be generated with the video.
  • Flexible framing: both 16:9 landscape and 9:16 portrait formats support conventional video and mobile-first content.
  • Image and frame control: supported image inputs and first-and-last-frame guidance can provide more structure than a text-only prompt.

The trade-off is that Fast is not the obvious choice for every final render. Its clips are limited to four, six, or eight seconds, endpoint capabilities vary, and standard Veo 3.1 may be preferable when maximum visual fidelity is more important than iteration speed or cost. The supplied research does not establish a numerical quality difference, so the choice should be validated with representative prompts and outputs.

Limitations to plan for

Veo 3.1 Fast has several constraints that affect production design:

  • It generates short clips rather than long-form video, with durations limited to four, six, or eight seconds.
  • English is the documented prompt language. Other languages may work inconsistently.
  • Generation latency can range from seconds to several minutes during periods of high demand, so “Fast” does not mean a guaranteed fixed response time.
  • Gemini API videos are retained on the server for a limited period and should be downloaded promptly.
  • Multi-video prompting and reasoning across multiple source videos are not supported.
  • Reference-image and video-extension features vary by endpoint and should not be assumed from the family name alone.
  • Safety filters, memorization checks, and audio-processing checks can block a requested result.
  • Generated video is watermarked with SynthID.

These limitations make the model better suited to modular workflows that generate, review, download, and assemble short clips than to a single request for a complete long-form production.

Reasoning, coding, and tool support

Veo 3.1 Fast is a video-generation model, not a reasoning or coding model. The supplied model data does not assign a reasoning score or coding score, and those capabilities are not applicable to its primary task. It accepts prompts that describe a desired scene, but that should not be confused with general-purpose reasoning across documents, software repositories, or complex multi-step questions.

The model data also lists no tool or function-calling support, no streaming output, no fine-tuning, no caching, and no batch API capability. Its structured output and JSON mode are not supported as model capabilities. Applications can still use surrounding application code to validate requests, manage files, track jobs, or assemble generated clips, but those are application-level functions rather than tools provided by Veo 3.1 Fast itself.

When to choose Veo 3.1 Fast

Choose Veo 3.1 Fast when the workflow benefits from producing many short audiovisual drafts quickly and economically. It is a practical option for:

  • Advertising teams testing multiple concepts, hooks, or visual treatments.
  • Social-media workflows that need portrait clips as well as standard landscape video.
  • Storyboarding and previsualization before committing to higher-quality renders.
  • Product demonstrations and narrated concept videos.
  • Automated applications that generate short video responses or content variations.
  • Creative teams that need native dialogue, effects, or ambience in early iterations.

Use standard Veo 3.1 instead when a final shot's visual fidelity is the dominant requirement and the higher-cost, slower iteration trade-off is acceptable. Consider Veo 3.1 Lite for especially cost-sensitive, high-volume work if its available features meet the project requirements. For long-form video, audio-only generation, general reasoning, or software development, a different model category is more appropriate.

Bottom line

Veo 3.1 Fast is best understood as a production-oriented iteration model: it creates short video with native audio, supports text and image guidance, and charges by generated second. Its value comes from the balance between speed, cost, and audiovisual output rather than from serving as the highest-fidelity or most general Veo option. Before deployment, confirm the exact endpoint's resolution, frame-control, video-input, and extension support, then test the model with the prompts and content policies that matter to the intended workflow.


Answers to Frequently Asked Questions

What are the Veo 3.1 Fast model identifiers for Gemini API and Vertex AI?
The Gemini API uses the preview model identifier `veo-3.1-fast-generate-preview`. Vertex AI uses `veo-3.1-fast-generate-001` for its production endpoint. Capabilities and availability can differ between the two services, so developers should verify the selected endpoint and region.
How much does Veo 3.1 Fast cost?
The Gemini API lists Veo 3.1 Fast at $0.10 per generated video second for 720p, $0.12 for 1080p, and $0.30 for 4K where supported. For example, a four-second 720p clip costs $0.40, while an eight-second 1080p clip costs $0.96.
What video formats and durations does Veo 3.1 Fast support?
Veo 3.1 Fast supports 16:9 landscape and 9:16 portrait video, with documented durations of four, six, or eight seconds at 24 frames per second. Supported inputs include text and images, while first-frame and last-frame guidance may be available depending on the endpoint.
What is Veo 3.1 Fast used for?
Veo 3.1 Fast is designed for generating short audiovisual clips quickly and at lower cost. It is useful for advertising concept testing, social-media variations, storyboarding, previsualization, product demonstrations, and other high-volume video workflows.
Does Veo 3.1 Fast generate audio?
Yes. Audio is always enabled, and prompts can request spoken dialogue, sound effects, and environmental ambience synchronized with the generated video. Outputs are watermarked with SynthID.


Sources 7
Provider

About Google DeepMind