Gemini 3.1 Flash Lite Image

Nano Banana 2 Lite

by Google DeepMind · Generally available; scheduled retirement June 28, 2027 or later

Nano Banana 2 Lite, also called Gemini 3.1 Flash Lite Image, is a generally available Google image model optimized for fast and inexpensive 1K generation and editing. It accepts text, image, and video context, supports up to 14 reference images, and can return text and images together. Its main trade-offs are 1K-only output and the lack of structured output, grounding, function calling, tuning, and other advanced application features.

Text Image generation Reasoning Coding
Nano Banana 2 Lite is the developer-facing name for Google's Gemini 3.1 Flash Lite Image model. Available through the Gemini API and Google Cloud's Gemini Enterprise Agent Platform, it prioritizes low latency, low operating cost, and rapid 1K image generation. It can create images from text, edit supplied images, and use video as visual context, but it does not provide the higher-resolution output or broader advanced features associated with larger image models.
Outputs

What Nano Banana 2 Lite can produce

Text Image generation
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Prompt caching Batch API Multimodal output
Model profile

Performance characteristics

3/10 Reasoning
1/10 Coding
10/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Gemini 3.1 Flash Lite Image
Model type Multimodal
Context window 66K tokens
Maximum output 4K tokens
Release date June 23, 2026
Status Generally available; scheduled retirement June 28, 2027 or later
Shutdown date June 28, 2027
Knowledge cutoff notes

No exact knowledge cutoff was identified in the reviewed first-party model, pricing, or Google Cloud documentation.

Model notes

Nano Banana 2 Lite is the product name for Gemini 3.1 Flash Lite Image, model ID gemini-3.1-flash-lite-image. It supports text-to-image generation, image editing, interleaved text-and-image responses, up to 14 reference images, video-to-image context, and 1K output only. Google documents sub-2-second target latency and positions it as the fastest and cheapest Gemini image model. Google Cloud lists implicit context caching, batch inference, and no structured output, grounding, function calling, tuning, Gemini Live API, or URL context. Generated images include SynthID watermarking and the model supports C2PA content credentials. The exact knowledge cutoff is not publicly specified in the reviewed first-party documentation.

Cost

Model pricing

Input $0.25 per 1 million input tokens for text, image, and video input; batch pricing $0.125 per 1 million input tokens
Output $1.50 per 1 million text/thought output tokens and $30 per 1 million image output tokens; approximately $0.0336 per 1K image; batch image output pricing $15 per 1 million image tokens
Model guide

Nano Banana 2 Lite: Fast, Low-Cost 1K Image Generation

Nano Banana 2 Lite is Google DeepMind's efficiency-focused Gemini image model for fast, inexpensive 1K image generation and editing. It accepts text, images, and video as context, supports interleaved text-and-image responses, and is designed for interactive and high-volume applications rather than high-resolution production work.

What is Nano Banana 2 Lite?

Nano Banana 2 Lite is Google DeepMind's efficiency-focused image-generation and image-editing model. Its official Gemini API model ID is gemini-3.1-flash-lite-image, and Google also documents it as Gemini 3.1 Flash Lite Image. “Nano Banana 2 Lite” is the product name used for this model in the supplied model information.

The model is built for situations where an application needs images quickly and at a relatively low per-image cost. It can generate an image from a text prompt, modify an image supplied by the user, and perform iterative edits over multiple turns. It can also use video as visual context when producing an image. This makes it more than a simple text-to-image endpoint, although its design is intentionally narrower than that of a larger, higher-resolution image model.

Unlike a text-only model that describes an image, Nano Banana 2 Lite can return generated image data directly. It also supports interleaved responses, meaning a response can contain both explanatory text and generated images. The model is therefore suited to applications where the user wants to request a visual result and receive a short explanation, label, or follow-up text in the same interaction.

Where it fits in Google's model lineup

Nano Banana 2 Lite occupies the speed-and-cost-oriented position in Google's current Nano Banana 2 family. Google positions it as the fastest and cheapest Gemini image model, with a focus on low latency and 1K output. That positioning involves clear trade-offs: the model is not intended to replace a higher-resolution model when the final asset must be produced at 2K or 4K, or when a workflow depends heavily on complex multi-reference control and sequential editing.

The larger Nano Banana 2 model is the more appropriate comparison when an application needs higher-resolution production output or more advanced reference-image workflows. Nano Banana Pro is another option for users evaluating Google's broader image-model range. Those models are mentioned here only as positioning references; Nano Banana 2 Lite is specifically optimized for lightweight, rapid image work.

Main capabilities and supported inputs

Nano Banana 2 Lite accepts multiple kinds of visual context. Its documented input and output capabilities include:

  • Text input: Text prompts can describe a new image or specify changes to an existing image.
  • Image input: The model can edit supplied images and accept up to 14 reference images in one prompt.
  • Video input: Video can provide visual context for image generation, with the video sampled at one frame per second.
  • Image output: The model generates 1K images and supports multiple aspect ratios.
  • Text output: It can return text alongside image output, enabling interleaved text-and-image responses.

Supported aspect ratios include common square, landscape, portrait, ultrawide, and tall formats. The documented examples include 1:1, 16:9, 9:16, 21:9, 1:4, and 4:1. This range is useful when the same application needs to produce social, presentation, banner, or mobile-oriented compositions without forcing every result into a square canvas.

For provenance, the supplied documentation states that generated images include Google's SynthID watermark and that the model supports C2PA content credentials. These mechanisms can help identify or attach provenance information to generated media, although they do not make the content automatically accurate or suitable for every publishing workflow.

Technical specifications and limits

SpecificationDocumented detail
Product nameNano Banana 2 Lite
Official model nameGemini 3.1 Flash Lite Image
API model IDgemini-3.1-flash-lite-image
Context window65,536 input tokens
Maximum output4,096 output tokens
Image resolution1K output only
Reference-image limitUp to 14 images per prompt
Maximum input size500 MB
Release stageGenerally available
Release dateJune 23, 2026
Scheduled shutdownJune 28, 2027, according to the supplied model record

The context window is the amount of tokenized input the model can process, while the 4,096-token figure describes the maximum output-token setting documented for the model. Neither figure means that an image itself contains that many tokens in a directly visible way. For practical users, the more important image-specific restriction is the 1K output ceiling.

The 500 MB maximum input size gives applications room to provide substantial visual context, but a larger upload does not necessarily produce a better result. Nano Banana 2 Lite is specifically described as less optimized for complex multiple-reference inputs and multi-turn sequential editing than the larger Nano Banana 2 model. Workflows that require precise control across many source images should therefore be tested carefully rather than assumed to scale linearly with the stated input limit.

Pricing and API availability

The supplied Gemini API pricing lists Nano Banana 2 Lite at $0.25 per one million input tokens for text, image, and video input. Text and thought output costs $1.50 per one million tokens. Generated image output costs $30 per one million image tokens.

Google's documented estimate for a typical 1K generated image is approximately 1,120 output image tokens. At the standard image-output rate, that works out to roughly $0.0336 per generated image, before accounting for input and any accompanying text output. Actual usage can vary with the response and the amount of input context.

Batch pricing is lower in the supplied pricing information: input is listed at $0.125 per one million tokens, text and thought output at $0.75 per one million tokens, and image output at $15 per one million image tokens. Batch processing is more relevant to queued or asynchronous workloads than to an interactive application that must respond immediately.

The model is available through the Gemini API and Google Cloud's Gemini Enterprise Agent Platform. The supplied cloud notes list implicit context caching and batch inference as supported. They also state that structured output, grounding, function calling, tuning, Gemini Live API access, and URL context are not available for this model. These restrictions matter when choosing an image model for a larger application architecture: Nano Banana 2 Lite can generate visual content, but it is not a general-purpose tool-orchestration model.

Speed, cost, and quality trade-offs

Nano Banana 2 Lite's principal advantage is operational efficiency. Google positions it for low latency, and the supplied notes describe a sub-two-second target latency. That makes it a practical candidate for an interactive editor, a consumer application that generates many visual variations, or a workflow where users repeatedly adjust a prompt and inspect the result.

The lower cost is particularly relevant when an application generates many images rather than a small number of final assets. Examples include thumbnail variations, advertising concepts, classroom illustrations, storyboards, product-color explorations, and user-generated image transformations. At approximately 3.36 cents in image-output cost for a typical 1K image at the standard rate, repeated experimentation can be less expensive than using a higher-end model for every draft.

These advantages are not free of trade-offs. The model produces 1K output only, so a team producing print-ready or high-resolution campaign assets may need a different model or an additional enlargement workflow. The supplied research also characterizes Nano Banana 2 Lite as less suitable for advanced multi-reference work, high-end sequential editing, and web-grounded image generation. Its speed and price should therefore be evaluated against the quality and control requirements of the complete workflow, not just the cost of a single request.

Reasoning, coding, and tool support

Nano Banana 2 Lite is primarily an image model, not a reasoning or coding model. It can interpret text instructions and visual context, but the supplied evaluation information gives it a low coding orientation and does not position it for software-development tasks. It should not be selected as the main model for code generation, repository analysis, complex planning, or tool-driven automation.

The documentation supplied for this model does not list function calling, grounding, URL context, or structured JSON output as supported features. That means an application requiring machine-readable structured responses, external search grounding, or direct function execution should place those responsibilities in surrounding application logic or choose a different model. Nano Banana 2 Lite can still be used as the visual-generation component of a larger pipeline, provided another component handles orchestration and validation.

Its reasoning behavior should also be interpreted narrowly. The model can follow image instructions and use supplied visual information, but there is no supplied benchmark establishing advanced reasoning performance. Any assessment of its ability to understand complicated scenes, preserve exact details, or follow long multi-step instructions should be treated as an application-testing question rather than a verified provider specification.

Best use cases

Nano Banana 2 Lite is a strong fit when the priority is fast visual iteration at controlled cost. Suitable applications include:

  • Rapid text-to-image prototyping and concept exploration.
  • High-volume generation of image variations for marketing or product testing.
  • Lightweight edits such as background changes, color adjustments, and local modifications.
  • Interactive consumer applications where users expect quick visual feedback.
  • Storyboarding, educational illustrations, and early-stage creative development.
  • Applications that need video frames or multiple reference images as visual context but do not require the highest-resolution final image.
  • Draft images for later review, selection, or processing by another system.

For example, a design tool could use Nano Banana 2 Lite to create several 1K background variations while a user changes a color or aspect ratio. A learning application could generate a simple illustration from a lesson description. A commerce workflow could produce multiple visual concepts for internal review. In each case, the benefit comes from fast iteration rather than maximum production fidelity.

When to choose Nano Banana 2 Lite

Choose Nano Banana 2 Lite when response speed, low operating cost, and 1K image generation matter more than maximum resolution. It is especially appropriate when users will generate many drafts, when image editing is relatively lightweight, or when the application needs to combine text, image, and video context in an interactive flow.

Consider the larger Nano Banana 2 model when the workflow needs 2K or 4K output, stronger handling of multiple reference images, or more advanced multi-turn sequential editing. Consider another model or a separate orchestration layer when the application depends on web grounding, function calling, structured JSON, tuning, or coding support. Nano Banana 2 Lite can still serve as the image-producing part of such a system, but the supplied documentation does not support treating it as a complete general-purpose agent.

Limitations and bottom line

The central limitation is the 1K-only output ceiling. Other important constraints include the lack of documented structured output, grounding, function calling, tuning, URL context, and Gemini Live API support. The model is also not positioned for coding tasks or demanding multi-reference and sequential-editing workflows.

Within those boundaries, Nano Banana 2 Lite has a clear purpose: fast, inexpensive image generation and editing for interactive or high-volume use. Its combination of text, image, and video context; interleaved text-and-image responses; broad aspect-ratio support; and low estimated per-image cost makes it useful for drafts, variations, lightweight edits, and visual experimentation. Teams that need final high-resolution assets or more elaborate image control should evaluate a larger sibling model instead of treating the Lite model as a universal replacement.


Answers to Frequently Asked Questions

When should I choose Nano Banana 2 Lite instead of a larger image model?
Choose Nano Banana 2 Lite when low latency, low cost, and rapid 1K image generation matter most. It is well suited to drafts, image variations, lightweight edits, storyboards, interactive applications, and high-volume experimentation. Choose a larger model when you need 2K or 4K output, advanced reference-image control, or more sophisticated sequential editing.
How many reference images and what types of input does Nano Banana 2 Lite support?
The model accepts text, image, and video input. It can use up to 14 reference images in one prompt, and video can provide visual context with frames sampled at one frame per second. The maximum input size is 500 MB.
What are the main limitations of Nano Banana 2 Lite?
Nano Banana 2 Lite produces 1K images only and is less suited to complex multi-reference workflows and advanced sequential editing. The supplied documentation also lists structured output, grounding, function calling, tuning, URL context, and Gemini Live API access as unsupported.
What is Nano Banana 2 Lite?
Nano Banana 2 Lite is Google DeepMind's efficiency-focused image-generation and image-editing model. Its official Gemini API model ID is gemini-3.1-flash-lite-image, also documented as Gemini 3.1 Flash Lite Image. It generates and edits 1K images using text, image, and video context.
How much does Nano Banana 2 Lite cost per generated image?
Image output is priced at $30 per one million image tokens. Google estimates that a typical 1K image uses about 1,120 output image tokens, resulting in an image-output cost of approximately $0.0336 per image, before input and accompanying text-output costs.


Sources 6
Provider

About Google DeepMind