What is Nano Banana?
Nano Banana is the former name and product alias for Gemini 2.5 Flash Image, Google DeepMind’s native image-generation and editing model. Its canonical Gemini API model identifier is gemini-2.5-flash-image.
The model is multimodal in both a practical and technical sense: it can accept text and images as input, understand the relationship between them, and produce text and generated images as output. For example, a user can provide a photograph and ask the model to change the background, alter an object, restyle the scene, or create a variation. The interaction can continue conversationally, allowing successive edits instead of requiring every instruction to be written as a completely new prompt.
Nano Banana is not intended to be a general-purpose reasoning model with broad tool access. Its primary role is visual creation and transformation where response speed, straightforward prompting, and image-generation throughput matter more than advanced reasoning or complex application control.
Where Nano Banana fits in Google’s model lineup
Nano Banana represents the Gemini 2.5 generation of Google’s image-capable models. It was introduced on August 26, 2025, and the stable model became generally available on October 2, 2025. Google has since deprecated it and announced that access will end on October 2, 2026.
That status changes how the model should be evaluated. It can still be useful for existing applications that depend on gemini-2.5-flash-image, but it is not the best foundation for a new long-lived integration. Google recommends migrating to Gemini 3.1 Flash Image, also known as Nano Banana 2, or Gemini 3.1 Flash Lite Image, also known as Nano Banana 2 Lite. Nano Banana Pro, based on Gemini 3 Pro Image, is positioned for higher-quality and more complex visual production.
These newer names are included only to explain Nano Banana’s current position. Nano Banana itself remains the subject: it is the older, speed-oriented image model that preceded Google’s newer Nano Banana variants.
What Nano Banana can do
Nano Banana combines image understanding with native image generation. Instead of treating image creation as a separate process from image analysis, it can use a supplied image as part of a text-guided visual conversation.
- Text-to-image generation: create an image from a written description.
- Conversational image editing: provide an image and describe changes in natural language, then refine the result through follow-up instructions.
- Image-to-image transformation: modify or restyle an existing image while using it as a visual reference.
- Visual interpretation: process image input alongside text instructions.
- Text and image output: return explanatory text as well as generated images.
- Resolution up to 1024 by 1024 pixels: the supplied research identifies this as the model’s supported image-generation ceiling.
These capabilities make the model suitable for tasks such as producing concept variations, adapting supplied product or marketing imagery, experimenting with visual styles, and building applications where a user repeatedly adjusts an image through short instructions.
Technical specifications and supported modalities
The following details describe the exact Gemini 2.5 Flash Image model rather than the wider Gemini product ecosystem.
| Specification | Nano Banana |
|---|---|
| Canonical model ID | gemini-2.5-flash-image |
| Input types | Text and images |
| Output types | Text and images |
| Input token limit | 65,536 tokens |
| Maximum output tokens | 32,768 tokens |
| Maximum listed image resolution | 1024 × 1024 pixels |
| Knowledge cutoff | June 2025 |
| Model status | Deprecated; scheduled to shut down October 2, 2026 |
Tokens are units used to measure text and other model input or output data; the input and output limits are not the same as an image’s pixel dimensions. The 65,536-token input limit describes how much contextual material the model can process in a request, while the 32,768-token maximum concerns generated model output. The supplied documentation does not establish that every request will use or need the maximum output allowance.
Nano Banana supports image input and image output, but it does not support audio or video input, audio output, video output, music output, embeddings, or speech output. It is therefore an image-focused multimodal model, not a general media model.
Nano Banana pricing
Google’s standard paid pricing is based on both tokens and generated images:
- Text and image input: $0.30 per 1 million input tokens.
- Generated image output: $0.039 per generated image.
- Batch and Flex input: $0.15 per 1 million input tokens.
- Batch and Flex image output: $0.0195 per generated image.
- Priority input: $0.54 per 1 million input tokens.
- Priority image output: $0.0702 per generated image.
The per-image charge is important when estimating costs. A request can have modest text and image input but still incur image-output charges for each generated image. Conversely, workflows that process large amounts of prompt or image data can be affected by token pricing as well.
Batch and Flex processing offer lower listed prices when the workload can use those processing modes. Priority processing costs more and is intended for cases where a higher-priority service tier justifies the premium. These prices are API usage rates, not a consumer subscription price, and the supplied research does not specify a recurring consumer plan for Nano Banana itself.
Speed, reasoning, and coding trade-offs
Nano Banana’s main strength is targeted visual throughput rather than broad reasoning. The supplied evaluation data gives it a high editorial speed score of 8 out of 10, a moderate cost score of 6 out of 10, a low reasoning score of 3 out of 10, and a low coding score of 1 out of 10. These scores are editorial assessments, not provider-published benchmarks.
In practical terms, the model is a sensible fit when the application needs an image quickly and the task can be expressed through visual instructions. It is less suitable when the request depends on extended logical analysis, sophisticated planning, software development, or reliable tool orchestration.
Nano Banana does not support function calling, code execution, search grounding, URL context, or the Live API. Function calling would allow a model to request actions from external application tools; because that capability is unavailable here, an application must handle external operations separately. The model also does not support structured outputs, so developers should not treat it as a suitable endpoint for guaranteed JSON-shaped responses. Caching is listed as unsupported as well.
Main strengths and limitations
Strengths
- Native image generation and editing: image creation and image transformation are central model capabilities rather than secondary add-ons.
- Conversational workflow: users can refine an image through successive natural-language instructions.
- Fast visual processing: the model is positioned for low-latency and high-volume image workflows.
- Text-and-image interaction: it can combine written instructions with supplied visual material.
- Simple API positioning: the canonical model identifier makes it clear which endpoint an existing application is using.
Limitations
- Deprecation: scheduled shutdown makes it a poor choice for a new application that needs long-term continuity.
- No audio or video support: it cannot handle audio or video workflows.
- No external tools: function calling, search grounding, code execution, URL context, and Live API access are unavailable.
- No structured outputs: applications needing dependable machine-readable JSON should use a different model or add their own handling layer.
- Limited reasoning and coding suitability: the model is optimized for visual tasks, not complex analysis or software engineering.
- Image-output costs: each generated image carries a separate listed charge in addition to token usage.
Best use cases for Nano Banana
Nano Banana is most appropriate when visual iteration is more important than advanced reasoning or long-term model availability. Suitable applications include:
- Rapid text-to-image prototyping.
- Conversational photo and illustration editing.
- Transforming a supplied image according to natural-language instructions.
- Generating multiple visual concepts for early design exploration.
- High-volume image workflows where latency and per-image cost matter.
- Creative tools that let users refine an image through short follow-up prompts.
For example, a design tool could let a user upload a room photograph and ask for lighter walls, different furniture, or a new visual style. A creative application could also generate several concept images from a short description, then use follow-up instructions to adjust composition or appearance.
When to choose Nano Banana—and when not to
Choose Nano Banana when you are maintaining an existing integration, need the specific behavior of Gemini 2.5 Flash Image, or require fast image generation and editing without tool calls. Its lower-cost Batch and Flex rates may also be useful for workloads that do not require priority processing.
For a new production integration, the scheduled shutdown is the decisive concern. A newer Nano Banana model is a more appropriate starting point when the application needs a supported future endpoint. Nano Banana 2 or Nano Banana 2 Lite should be considered when a newer image-generation model is preferred, while Nano Banana Pro is the better direction for higher-quality or more complex visual production according to Google’s positioning.
A different type of model is more appropriate if the application needs audio or video handling, structured JSON, external tools, web search, code execution, or sophisticated reasoning. Nano Banana can contribute to a larger pipeline, but the supplied specifications do not support treating it as the central general-purpose model for those tasks.
Bottom line
Nano Banana is a focused, fast image-generation and editing model rather than a general-purpose AI assistant. Its useful combination of text and image input, conversational editing, image output, and relatively simple pricing makes it practical for visual experimentation and existing high-volume workflows. However, its lack of tool support, structured outputs, audio and video capabilities, and advanced reasoning features limits its scope. Because Google has deprecated the model and scheduled its shutdown for October 2, 2026, new projects should generally evaluate its successors instead of building around gemini-2.5-flash-image.

