What is Nano Banana 2 Lite?
Nano Banana 2 Lite is Google DeepMind's efficiency-focused image-generation and image-editing model. Its official Gemini API model ID is gemini-3.1-flash-lite-image, and Google also documents it as Gemini 3.1 Flash Lite Image. “Nano Banana 2 Lite” is the product name used for this model in the supplied model information.
The model is built for situations where an application needs images quickly and at a relatively low per-image cost. It can generate an image from a text prompt, modify an image supplied by the user, and perform iterative edits over multiple turns. It can also use video as visual context when producing an image. This makes it more than a simple text-to-image endpoint, although its design is intentionally narrower than that of a larger, higher-resolution image model.
Unlike a text-only model that describes an image, Nano Banana 2 Lite can return generated image data directly. It also supports interleaved responses, meaning a response can contain both explanatory text and generated images. The model is therefore suited to applications where the user wants to request a visual result and receive a short explanation, label, or follow-up text in the same interaction.
Where it fits in Google's model lineup
Nano Banana 2 Lite occupies the speed-and-cost-oriented position in Google's current Nano Banana 2 family. Google positions it as the fastest and cheapest Gemini image model, with a focus on low latency and 1K output. That positioning involves clear trade-offs: the model is not intended to replace a higher-resolution model when the final asset must be produced at 2K or 4K, or when a workflow depends heavily on complex multi-reference control and sequential editing.
The larger Nano Banana 2 model is the more appropriate comparison when an application needs higher-resolution production output or more advanced reference-image workflows. Nano Banana Pro is another option for users evaluating Google's broader image-model range. Those models are mentioned here only as positioning references; Nano Banana 2 Lite is specifically optimized for lightweight, rapid image work.
Main capabilities and supported inputs
Nano Banana 2 Lite accepts multiple kinds of visual context. Its documented input and output capabilities include:
- Text input: Text prompts can describe a new image or specify changes to an existing image.
- Image input: The model can edit supplied images and accept up to 14 reference images in one prompt.
- Video input: Video can provide visual context for image generation, with the video sampled at one frame per second.
- Image output: The model generates 1K images and supports multiple aspect ratios.
- Text output: It can return text alongside image output, enabling interleaved text-and-image responses.
Supported aspect ratios include common square, landscape, portrait, ultrawide, and tall formats. The documented examples include 1:1, 16:9, 9:16, 21:9, 1:4, and 4:1. This range is useful when the same application needs to produce social, presentation, banner, or mobile-oriented compositions without forcing every result into a square canvas.
For provenance, the supplied documentation states that generated images include Google's SynthID watermark and that the model supports C2PA content credentials. These mechanisms can help identify or attach provenance information to generated media, although they do not make the content automatically accurate or suitable for every publishing workflow.
Technical specifications and limits
| Specification | Documented detail |
|---|---|
| Product name | Nano Banana 2 Lite |
| Official model name | Gemini 3.1 Flash Lite Image |
| API model ID | gemini-3.1-flash-lite-image |
| Context window | 65,536 input tokens |
| Maximum output | 4,096 output tokens |
| Image resolution | 1K output only |
| Reference-image limit | Up to 14 images per prompt |
| Maximum input size | 500 MB |
| Release stage | Generally available |
| Release date | June 23, 2026 |
| Scheduled shutdown | June 28, 2027, according to the supplied model record |
The context window is the amount of tokenized input the model can process, while the 4,096-token figure describes the maximum output-token setting documented for the model. Neither figure means that an image itself contains that many tokens in a directly visible way. For practical users, the more important image-specific restriction is the 1K output ceiling.
The 500 MB maximum input size gives applications room to provide substantial visual context, but a larger upload does not necessarily produce a better result. Nano Banana 2 Lite is specifically described as less optimized for complex multiple-reference inputs and multi-turn sequential editing than the larger Nano Banana 2 model. Workflows that require precise control across many source images should therefore be tested carefully rather than assumed to scale linearly with the stated input limit.
Pricing and API availability
The supplied Gemini API pricing lists Nano Banana 2 Lite at $0.25 per one million input tokens for text, image, and video input. Text and thought output costs $1.50 per one million tokens. Generated image output costs $30 per one million image tokens.
Google's documented estimate for a typical 1K generated image is approximately 1,120 output image tokens. At the standard image-output rate, that works out to roughly $0.0336 per generated image, before accounting for input and any accompanying text output. Actual usage can vary with the response and the amount of input context.
Batch pricing is lower in the supplied pricing information: input is listed at $0.125 per one million tokens, text and thought output at $0.75 per one million tokens, and image output at $15 per one million image tokens. Batch processing is more relevant to queued or asynchronous workloads than to an interactive application that must respond immediately.
The model is available through the Gemini API and Google Cloud's Gemini Enterprise Agent Platform. The supplied cloud notes list implicit context caching and batch inference as supported. They also state that structured output, grounding, function calling, tuning, Gemini Live API access, and URL context are not available for this model. These restrictions matter when choosing an image model for a larger application architecture: Nano Banana 2 Lite can generate visual content, but it is not a general-purpose tool-orchestration model.
Speed, cost, and quality trade-offs
Nano Banana 2 Lite's principal advantage is operational efficiency. Google positions it for low latency, and the supplied notes describe a sub-two-second target latency. That makes it a practical candidate for an interactive editor, a consumer application that generates many visual variations, or a workflow where users repeatedly adjust a prompt and inspect the result.
The lower cost is particularly relevant when an application generates many images rather than a small number of final assets. Examples include thumbnail variations, advertising concepts, classroom illustrations, storyboards, product-color explorations, and user-generated image transformations. At approximately 3.36 cents in image-output cost for a typical 1K image at the standard rate, repeated experimentation can be less expensive than using a higher-end model for every draft.
These advantages are not free of trade-offs. The model produces 1K output only, so a team producing print-ready or high-resolution campaign assets may need a different model or an additional enlargement workflow. The supplied research also characterizes Nano Banana 2 Lite as less suitable for advanced multi-reference work, high-end sequential editing, and web-grounded image generation. Its speed and price should therefore be evaluated against the quality and control requirements of the complete workflow, not just the cost of a single request.
Reasoning, coding, and tool support
Nano Banana 2 Lite is primarily an image model, not a reasoning or coding model. It can interpret text instructions and visual context, but the supplied evaluation information gives it a low coding orientation and does not position it for software-development tasks. It should not be selected as the main model for code generation, repository analysis, complex planning, or tool-driven automation.
The documentation supplied for this model does not list function calling, grounding, URL context, or structured JSON output as supported features. That means an application requiring machine-readable structured responses, external search grounding, or direct function execution should place those responsibilities in surrounding application logic or choose a different model. Nano Banana 2 Lite can still be used as the visual-generation component of a larger pipeline, provided another component handles orchestration and validation.
Its reasoning behavior should also be interpreted narrowly. The model can follow image instructions and use supplied visual information, but there is no supplied benchmark establishing advanced reasoning performance. Any assessment of its ability to understand complicated scenes, preserve exact details, or follow long multi-step instructions should be treated as an application-testing question rather than a verified provider specification.
Best use cases
Nano Banana 2 Lite is a strong fit when the priority is fast visual iteration at controlled cost. Suitable applications include:
- Rapid text-to-image prototyping and concept exploration.
- High-volume generation of image variations for marketing or product testing.
- Lightweight edits such as background changes, color adjustments, and local modifications.
- Interactive consumer applications where users expect quick visual feedback.
- Storyboarding, educational illustrations, and early-stage creative development.
- Applications that need video frames or multiple reference images as visual context but do not require the highest-resolution final image.
- Draft images for later review, selection, or processing by another system.
For example, a design tool could use Nano Banana 2 Lite to create several 1K background variations while a user changes a color or aspect ratio. A learning application could generate a simple illustration from a lesson description. A commerce workflow could produce multiple visual concepts for internal review. In each case, the benefit comes from fast iteration rather than maximum production fidelity.
When to choose Nano Banana 2 Lite
Choose Nano Banana 2 Lite when response speed, low operating cost, and 1K image generation matter more than maximum resolution. It is especially appropriate when users will generate many drafts, when image editing is relatively lightweight, or when the application needs to combine text, image, and video context in an interactive flow.
Consider the larger Nano Banana 2 model when the workflow needs 2K or 4K output, stronger handling of multiple reference images, or more advanced multi-turn sequential editing. Consider another model or a separate orchestration layer when the application depends on web grounding, function calling, structured JSON, tuning, or coding support. Nano Banana 2 Lite can still serve as the image-producing part of such a system, but the supplied documentation does not support treating it as a complete general-purpose agent.
Limitations and bottom line
The central limitation is the 1K-only output ceiling. Other important constraints include the lack of documented structured output, grounding, function calling, tuning, URL context, and Gemini Live API support. The model is also not positioned for coding tasks or demanding multi-reference and sequential-editing workflows.
Within those boundaries, Nano Banana 2 Lite has a clear purpose: fast, inexpensive image generation and editing for interactive or high-volume use. Its combination of text, image, and video context; interleaved text-and-image responses; broad aspect-ratio support; and low estimated per-image cost makes it useful for drafts, variations, lightweight edits, and visual experimentation. Teams that need final high-resolution assets or more elaborate image control should evaluate a larger sibling model instead of treating the Lite model as a universal replacement.
Answers to Frequently Asked Questions
gemini-3.1-flash-lite-image, also documented as Gemini 3.1 Flash Lite Image. It generates and edits 1K images using text, image, and video context.
