What is Gemini Omni Flash?
Gemini Omni Flash is a Google DeepMind multimodal model focused on video generation and editing. Its stable Gemini API identifier is gemini-omni-1.1-flash; Google also documents gemini-omni-flash-preview as a preview identifier. In practical terms, it is a specialized creative model: its main output is video, not a text response, code completion, image, or structured JSON object.
The model accepts combinations of text, images, and video for supported workflows. A user can describe a scene, provide a still image to animate, upload a short clip for modification, or use a previous interaction to continue refining a result. Generated clips can also include audio in supported generation scenarios.
Where it fits in Google's current lineup
Gemini Omni Flash belongs to Google's Gemini Omni family and occupies a video-first position within Google's current model catalog. Whereas a general Gemini model would normally be selected for conversation, coding, research, or document analysis, Omni Flash is intended for producing and manipulating short moving-image content.
The “Flash” designation is relevant to its positioning: the supplied research describes it as optimized for fast video generation and iterative creative work. The speed advantage is paired with a narrower scope. It is not presented as a general reasoning model, coding assistant, embedding model, transcription system, or standalone text-to-speech model.
What can the model do?
- Text-to-video: Create a short video from a natural-language description, such as a product demonstration or cinematic scene.
- Image-to-video: Animate a photograph, illustration, product image, or other still reference.
- Video editing: Apply conversational instructions to an uploaded or previously generated clip.
- Video extension: Continue a clip by appending newly generated content to its end.
- Frame interpolation: Generate a transition between a starting image and an ending image.
- Upscaling: Produce higher-resolution versions, including 1080p and 4K outputs, in addition to 360p and 720p generation.
- Conversational refinement: Revise a result over multiple turns using the Interactions API and a previous interaction ID.
These capabilities make the model useful for a workflow in which a creator starts with a rough concept, evaluates a short result, and then asks for targeted changes. The interaction model is particularly relevant when the desired result is difficult to specify perfectly in one prompt.
Technical specifications and supported modalities
| Specification | Documented detail |
|---|---|
| Stable model ID | gemini-omni-1.1-flash |
| Preview model ID | gemini-omni-flash-preview |
| Input modalities | Text, image, and video for the documented creative workflows |
| Primary output | Video; generated clips may include audio |
| Context window | 1,048,576 tokens |
| Typical generation length | Approximately 3 to 10 seconds per generation |
| Extended workflow length | Up to 40 seconds in supported video-extension workflows |
| Resolutions | 360p, 720p, 1080p, and 4K |
| Frame rate | 24 FPS |
| API | Gemini API Interactions API |
The large context window is a published specification, but it should not be interpreted as a guarantee that every input combination or video duration is accepted without other restrictions. Video-specific limits still apply. For example, uploaded videos used for editing or extension generally must be 10 seconds or shorter unless a supported multi-turn workflow applies.
Pricing and API delivery
Google's documented standard pricing for Gemini Omni Flash Preview is $1.50 per 1 million input tokens for text, image, video, and audio inputs. Output pricing is $9.00 per 1 million text tokens where applicable or $17.50 per 1 million video output tokens. Google calculates 720p video at 5,792 output tokens per second, which the supplied pricing information equates to approximately $0.10 per second under standard pricing.
These token prices are not the same as a fixed per-clip subscription price. The cost of a video request depends on the input and generated output represented in billable tokens, including the duration and resolution of video output. Pricing and model availability can change, so production users should verify the current Google pricing page before estimating a large batch workflow.
The Interactions API can return generated video inline or through a URI delivery mode. URI delivery is recommended for larger files. Streaming responses are supported through server-sent events, and previous interaction IDs enable conversational editing across turns.
Reasoning, coding, and tool support
Gemini Omni Flash should be evaluated as a video-generation model, not as a conventional reasoning or coding model. The supplied editorial assessment gives it a reasoning score of 5 out of 10 and a coding score of 1 out of 10. These are editorial estimates, not Google benchmark results or provider-published ratings.
Its reasoning ability is mainly useful for interpreting creative instructions, maintaining an intended visual direction, and applying revision requests. It is not documented here as a model for complex analytical answers, software engineering, embeddings, speech transcription, or standalone speech generation.
The supplied model data records no conventional tool-use or function-calling support and no structured-output capability. It also records text output as unsupported as a primary output type. Developers who need reliable JSON responses, function calls, or a text-first agent should choose a model designed for those tasks instead of adapting Omni Flash around a video workflow.
Strengths and trade-offs
The model's central strength is the combination of speed and video-specific iteration. A creator can move from a text idea or still image to a short clip, then revise the result using conversational instructions. Video extension, interpolation, and upscaling add useful operations that are not limited to one-shot text-to-video generation.
The main trade-off is specialization. The model's short clip duration, 24 FPS output, regional restrictions, and video-oriented API make it a better fit for rapid creative production than for long-form filmmaking or general AI automation. The editorial cost score is 6 out of 10 and speed score is 9 out of 10; these scores are subjective comparisons supplied for orientation, not vendor measurements.
Important limitations
- Short source and output clips: Individual generations are typically 3 to 10 seconds, and uploaded videos for editing or extension generally need to be 10 seconds or shorter. Supported workflows can extend generated content up to a total of 40 seconds.
- Extension direction: Video extension appends content to the end of a clip. It does not prepend footage or extend the middle of a sequence.
- Regional availability: Uploaded-video editing and extension have restrictions in regions including the European Economic Area, Switzerland, and the United Kingdom.
- Audio references: Standalone audio reference uploads are not supported in the current API. Audio contained in video references may be ignored, and voice editing is not supported.
- Limited generation controls: System instructions, temperature, top-p, stop sequences, and negative-prompt parameters are not available as separate controls in the documented API version.
- Provenance marking: Generated videos include Google's invisible SynthID watermarking for provenance verification.
The audio distinction matters in practice. Generated clips may contain audio, and pricing documentation lists audio among input categories, but that does not mean the model supports arbitrary standalone audio references or voice-directed video editing.
Best use cases
Gemini Omni Flash is a strong candidate for fast visual prototyping and short-form production tasks such as:
- Animating product photographs for advertisements or social posts
- Creating short marketing and promotional clips
- Turning illustrations or concept art into moving scenes
- Testing multiple visual directions before investing in a longer production
- Generating transitions between defined start and end images
- Extending a short shot with additional ending footage
- Iteratively revising a clip through conversational API interactions
- Producing short cinematic or visual-effects experiments
It is especially appropriate when speed and rapid variation matter more than long continuous duration, detailed low-level generation controls, or a conventional text response.
When to choose Gemini Omni Flash
Choose Gemini Omni Flash when the primary deliverable is a short video and you want one multimodal workflow for generation, editing, extension, interpolation, and upscaling. It is a sensible option for creative teams that prefer to refine results interactively rather than manage separate models for every video operation.
Consider another option when the task is primarily text chat, coding, structured JSON generation, function calling, transcription, standalone speech synthesis, or long-form video production. A general-purpose model may be more appropriate for planning, scripting, and software tasks, while a specialized production pipeline may offer better control for longer videos, detailed audio editing, or complex timeline work.
Overall, Gemini Omni Flash's value comes from its narrow but practical focus: fast, iterative video creation through a multimodal API. Its capabilities are substantial for short clips, but its limitations should be treated as design constraints rather than minor omissions.

