Gemini Omni

Gemini Omni Flash

by Google DeepMind · Current; stable model available as gemini-omni-1.1-flash, with gemini-omni-flash-preview also documented

Google DeepMind's Gemini Omni Flash is a video-first multimodal model for generating and editing short clips from text, images, and video. It supports conversational revisions through the Interactions API, video extension, frame interpolation, and outputs up to 4K, while imposing short-clip, regional, audio-reference, and control limitations.

Video generation Audio Reasoning Coding
Gemini Omni Flash is a Google DeepMind model built for short-form video creation rather than general-purpose text chat. It can generate clips from prompts and still images, edit existing videos, extend generated footage, interpolate between beginning and ending images, and support conversational revisions. The model is designed for users who value rapid creative iteration and video-specific controls, although its regional restrictions, short clip limits, unsupported audio-reference workflow, and lack of conventional structured-output features make it unsuitable for many non-video applications.
Outputs

What Gemini Omni Flash can produce

Video generation Audio
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

5/10 Reasoning
1/10 Coding
9/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Gemini Omni
Model type Multimodal
Context window 1.05M tokens
Release date June 30, 2026 (public preview); stable Gemini Omni 1.1 Flash availability documented in August 2026
Status Current; stable model available as gemini-omni-1.1-flash, with gemini-omni-flash-preview also documented
Knowledge cutoff notes

Google's current model documentation does not publish a distinct knowledge-cutoff date for Gemini Omni Flash.

Model notes

The canonical stable model identifier is gemini-omni-1.1-flash; gemini-omni-flash-preview is the preview identifier. The model generates 3–10 second clips at 360p, 720p, 1080p, or 4K and supports 24 FPS output. Generated videos can include audio. Google documents conversational editing through the Interactions API and previous interaction IDs. Video extension can append generated content up to a total of 40 seconds in supported workflows. Uploaded video editing and extension are regionally restricted, and standalone audio reference uploads are unsupported in the current API. The comparative scores are editorial estimates, not vendor benchmarks.

Cost

Model pricing

Input $1.50 per 1M input tokens for text, image, video, and audio inputs under standard pricing
Output $17.50 per 1M video output tokens; $9.00 per 1M text output tokens where applicable
Model guide

Gemini Omni Flash: Fast Multimodal Video Generation and Editing

Gemini Omni Flash is Google's fast multimodal video model for creating, extending, editing, interpolating, and upscaling short video clips from text, images, and video through the Gemini API Interactions API.

What is Gemini Omni Flash?

Gemini Omni Flash is a Google DeepMind multimodal model focused on video generation and editing. Its stable Gemini API identifier is gemini-omni-1.1-flash; Google also documents gemini-omni-flash-preview as a preview identifier. In practical terms, it is a specialized creative model: its main output is video, not a text response, code completion, image, or structured JSON object.

The model accepts combinations of text, images, and video for supported workflows. A user can describe a scene, provide a still image to animate, upload a short clip for modification, or use a previous interaction to continue refining a result. Generated clips can also include audio in supported generation scenarios.

Where it fits in Google's current lineup

Gemini Omni Flash belongs to Google's Gemini Omni family and occupies a video-first position within Google's current model catalog. Whereas a general Gemini model would normally be selected for conversation, coding, research, or document analysis, Omni Flash is intended for producing and manipulating short moving-image content.

The “Flash” designation is relevant to its positioning: the supplied research describes it as optimized for fast video generation and iterative creative work. The speed advantage is paired with a narrower scope. It is not presented as a general reasoning model, coding assistant, embedding model, transcription system, or standalone text-to-speech model.

What can the model do?

  • Text-to-video: Create a short video from a natural-language description, such as a product demonstration or cinematic scene.
  • Image-to-video: Animate a photograph, illustration, product image, or other still reference.
  • Video editing: Apply conversational instructions to an uploaded or previously generated clip.
  • Video extension: Continue a clip by appending newly generated content to its end.
  • Frame interpolation: Generate a transition between a starting image and an ending image.
  • Upscaling: Produce higher-resolution versions, including 1080p and 4K outputs, in addition to 360p and 720p generation.
  • Conversational refinement: Revise a result over multiple turns using the Interactions API and a previous interaction ID.

These capabilities make the model useful for a workflow in which a creator starts with a rough concept, evaluates a short result, and then asks for targeted changes. The interaction model is particularly relevant when the desired result is difficult to specify perfectly in one prompt.

Technical specifications and supported modalities

SpecificationDocumented detail
Stable model IDgemini-omni-1.1-flash
Preview model IDgemini-omni-flash-preview
Input modalitiesText, image, and video for the documented creative workflows
Primary outputVideo; generated clips may include audio
Context window1,048,576 tokens
Typical generation lengthApproximately 3 to 10 seconds per generation
Extended workflow lengthUp to 40 seconds in supported video-extension workflows
Resolutions360p, 720p, 1080p, and 4K
Frame rate24 FPS
APIGemini API Interactions API

The large context window is a published specification, but it should not be interpreted as a guarantee that every input combination or video duration is accepted without other restrictions. Video-specific limits still apply. For example, uploaded videos used for editing or extension generally must be 10 seconds or shorter unless a supported multi-turn workflow applies.

Pricing and API delivery

Google's documented standard pricing for Gemini Omni Flash Preview is $1.50 per 1 million input tokens for text, image, video, and audio inputs. Output pricing is $9.00 per 1 million text tokens where applicable or $17.50 per 1 million video output tokens. Google calculates 720p video at 5,792 output tokens per second, which the supplied pricing information equates to approximately $0.10 per second under standard pricing.

These token prices are not the same as a fixed per-clip subscription price. The cost of a video request depends on the input and generated output represented in billable tokens, including the duration and resolution of video output. Pricing and model availability can change, so production users should verify the current Google pricing page before estimating a large batch workflow.

The Interactions API can return generated video inline or through a URI delivery mode. URI delivery is recommended for larger files. Streaming responses are supported through server-sent events, and previous interaction IDs enable conversational editing across turns.

Reasoning, coding, and tool support

Gemini Omni Flash should be evaluated as a video-generation model, not as a conventional reasoning or coding model. The supplied editorial assessment gives it a reasoning score of 5 out of 10 and a coding score of 1 out of 10. These are editorial estimates, not Google benchmark results or provider-published ratings.

Its reasoning ability is mainly useful for interpreting creative instructions, maintaining an intended visual direction, and applying revision requests. It is not documented here as a model for complex analytical answers, software engineering, embeddings, speech transcription, or standalone speech generation.

The supplied model data records no conventional tool-use or function-calling support and no structured-output capability. It also records text output as unsupported as a primary output type. Developers who need reliable JSON responses, function calls, or a text-first agent should choose a model designed for those tasks instead of adapting Omni Flash around a video workflow.

Strengths and trade-offs

The model's central strength is the combination of speed and video-specific iteration. A creator can move from a text idea or still image to a short clip, then revise the result using conversational instructions. Video extension, interpolation, and upscaling add useful operations that are not limited to one-shot text-to-video generation.

The main trade-off is specialization. The model's short clip duration, 24 FPS output, regional restrictions, and video-oriented API make it a better fit for rapid creative production than for long-form filmmaking or general AI automation. The editorial cost score is 6 out of 10 and speed score is 9 out of 10; these scores are subjective comparisons supplied for orientation, not vendor measurements.

Important limitations

  • Short source and output clips: Individual generations are typically 3 to 10 seconds, and uploaded videos for editing or extension generally need to be 10 seconds or shorter. Supported workflows can extend generated content up to a total of 40 seconds.
  • Extension direction: Video extension appends content to the end of a clip. It does not prepend footage or extend the middle of a sequence.
  • Regional availability: Uploaded-video editing and extension have restrictions in regions including the European Economic Area, Switzerland, and the United Kingdom.
  • Audio references: Standalone audio reference uploads are not supported in the current API. Audio contained in video references may be ignored, and voice editing is not supported.
  • Limited generation controls: System instructions, temperature, top-p, stop sequences, and negative-prompt parameters are not available as separate controls in the documented API version.
  • Provenance marking: Generated videos include Google's invisible SynthID watermarking for provenance verification.

The audio distinction matters in practice. Generated clips may contain audio, and pricing documentation lists audio among input categories, but that does not mean the model supports arbitrary standalone audio references or voice-directed video editing.

Best use cases

Gemini Omni Flash is a strong candidate for fast visual prototyping and short-form production tasks such as:

  • Animating product photographs for advertisements or social posts
  • Creating short marketing and promotional clips
  • Turning illustrations or concept art into moving scenes
  • Testing multiple visual directions before investing in a longer production
  • Generating transitions between defined start and end images
  • Extending a short shot with additional ending footage
  • Iteratively revising a clip through conversational API interactions
  • Producing short cinematic or visual-effects experiments

It is especially appropriate when speed and rapid variation matter more than long continuous duration, detailed low-level generation controls, or a conventional text response.

When to choose Gemini Omni Flash

Choose Gemini Omni Flash when the primary deliverable is a short video and you want one multimodal workflow for generation, editing, extension, interpolation, and upscaling. It is a sensible option for creative teams that prefer to refine results interactively rather than manage separate models for every video operation.

Consider another option when the task is primarily text chat, coding, structured JSON generation, function calling, transcription, standalone speech synthesis, or long-form video production. A general-purpose model may be more appropriate for planning, scripting, and software tasks, while a specialized production pipeline may offer better control for longer videos, detailed audio editing, or complex timeline work.

Overall, Gemini Omni Flash's value comes from its narrow but practical focus: fast, iterative video creation through a multimodal API. Its capabilities are substantial for short clips, but its limitations should be treated as design constraints rather than minor omissions.


Answers to Frequently Asked Questions

What are the main limitations of Gemini Omni Flash?
Gemini Omni Flash is specialized for short video workflows rather than general reasoning, coding, structured JSON, function calling, transcription, or standalone speech generation. It cannot prepend footage or extend the middle of a sequence, standalone audio references and voice editing are unsupported, and uploaded-video editing or extension may be restricted in the European Economic Area, Switzerland, and the United Kingdom.
How much does Gemini Omni Flash cost?
For Gemini Omni Flash Preview, documented standard pricing is $1.50 per 1 million input tokens for text, image, video, and audio inputs, and $17.50 per 1 million video output tokens. Google calculates 720p video at 5,792 output tokens per second, which is approximately $0.10 per second under the supplied pricing information. Pricing and availability may change.
What video formats, durations, and resolutions does Gemini Omni Flash support?
Typical generations are approximately 3 to 10 seconds long at 24 FPS, with supported resolutions of 360p, 720p, 1080p, and 4K. Supported video-extension workflows can reach up to 40 seconds, while uploaded videos for editing or extension generally must be 10 seconds or shorter.
What is Gemini Omni Flash used for?
Gemini Omni Flash is a video-first multimodal model for generating and editing short videos. It supports text-to-video, image-to-video, conversational video editing, video extension, frame interpolation, and upscaling.
What are the model IDs for Gemini Omni Flash?
The stable model ID is "gemini-omni-1.1-flash". Google also documents "gemini-omni-flash-preview" as a preview identifier.


Sources 5
Provider

About Google DeepMind