What is Nano Banana 2?
Nano Banana 2 is Google DeepMind’s marketing name for Gemini 3.1 Flash Image, a native multimodal model designed primarily for image generation and image editing. Its stable Gemini API model identifier is gemini-3.1-flash-image.
Unlike a text-only language model, Nano Banana 2 can understand visual material and produce images as part of a conversational workflow. A user can provide a written instruction, one or more reference images, video, or a PDF, then ask the model to create, modify, translate, localize, upscale, or recompose visual content. Follow-up prompts can refine the result without restarting the entire task.
Within Google’s Nano Banana family, this model is positioned as the general-purpose, high-throughput option. It aims to retain many of the control and quality features associated with higher-end image generation while offering the speed expected from a Flash model. That positioning makes it more suitable for repeated visual iteration and production pipelines than for open-ended text reasoning or agentic automation.
What Nano Banana 2 can create and edit
The model supports both text-to-image generation and conversational editing. Practical requests can include creating a marketing poster, changing the style of a product photograph, translating text within an image, turning a video frame into a cinematic thumbnail, or combining information from a PDF with a visual design brief.
- New images generated from text prompts
- Edits based on one or more reference images
- Multi-turn refinement of an existing image
- Posters, social media graphics, mockups, diagrams, and infographics
- Image localization and translation
- Visual summaries based on video or PDF material
- Current or location-specific visuals using Google Search and Image Search grounding
Google documents support for 512-pixel, 1K, 2K, and 4K image outputs. The model also supports unusually wide and tall formats, including 1:4, 4:1, 1:8, and 8:1, alongside more conventional aspect ratios. These formats are useful for banners, mobile layouts, panoramic graphics, and other designs that do not fit a standard landscape or portrait frame.
Inputs, outputs, and visual understanding
Nano Banana 2 accepts text, images, video, and PDF files. Video can provide visual context for tasks such as selecting a scene for a thumbnail, designing a cinematic poster, or producing a summary infographic. Audio input is not supported for this image-generation model.
The model can return both generated images and text. The image is the important non-text output: accompanying text may describe the result or provide a response to the instruction, but Nano Banana 2 is not primarily intended as a conversational text model. It does not generate video, audio, music, speech, or embeddings.
Google also highlights improved multilingual text rendering. This is relevant when an image needs labels, headlines, signs, product copy, or other readable language. Text inside generated images can still require careful prompting and verification. Google notes that results may be better when the desired wording is specified precisely or developed before the rendering step rather than treated as an afterthought.
Grounding, consistency, and creative control
Nano Banana 2 can use Google Search and Google Image Search grounding. Grounding means that the model can consult current external information or visual references during a request, which is useful for subjects such as recent products, current events, locations, and factual visual explanations. Grounding is an API or platform feature; it does not change the model’s underlying knowledge cutoff, which is listed as January 2025 for this model line.
The model supports configurable thinking behavior, improved instruction following, stronger aspect-ratio adherence, and visual consistency across an iterative workflow. Google reports that it can preserve the resemblance of up to four characters and the fidelity of up to ten objects in a single workflow. Those figures are provider claims rather than an independent guarantee for every prompt or composition.
Generated images include a SynthID watermark. Google also documents a limitation on Image Search grounding: it does not currently support using real-world images of people from web search. As with other generative image systems, users should review outputs for factual accuracy, unwanted alterations, illegible text, and resemblance issues before publication.
Technical specifications at a glance
| Specification | Documented detail |
|---|---|
| Model name | Gemini 3.1 Flash Image |
| Marketing name | Nano Banana 2 |
| Stable model ID | gemini-3.1-flash-image |
| Context length | 131,072 tokens |
| Maximum output tokens | 32,768 |
| Input types | Text, image, video, and PDF |
| Output types | Image and text |
| Image resolutions | 512px, 1K, 2K, and 4K |
| Image grounding | Google Search and Google Image Search |
| Fine-tuning | Not supported according to the supplied model documentation |
| Caching | Not supported according to the supplied model documentation |
The context length describes the total amount of model context that can be supplied to a request, including relevant text and multimodal material as measured by the platform. It should not be interpreted as a promise that every large collection of files will produce equally strong results. The maximum output-token value is a platform specification for generated text and thinking output; image dimensions and image-output token accounting are separate concerns.
Reasoning, coding, and tool support
Nano Banana 2 includes configurable thinking behavior, which can help with complex visual instructions, composition decisions, and multi-step editing. However, it is still an image-focused model rather than a general reasoning model. The supplied editorial evaluation rates its reasoning at 7 out of 10 and coding at 2 out of 10; these are subjective catalog scores, not Google-published benchmark results.
The model documentation lists function calling, structured outputs, code execution, file search, URL context, Google Maps grounding, audio generation, and Live API support as unavailable. This distinction matters when selecting an API model. Nano Banana 2 can use Search or Image Search grounding for supported visual generation tasks, but it is not intended to operate as a tool-calling agent that plans and executes business actions.
For example, it may be appropriate to create a location-themed poster using grounded information, but not to build an assistant that reliably calls a booking system, returns strict JSON, executes code, or manages a multi-step workflow through external functions. A separate language or agent model is more appropriate for those jobs, with Nano Banana 2 used as the visual generation component when needed.
API pricing and availability
Nano Banana 2 is available through the Gemini API and Google AI Studio, with enterprise deployment through Google Cloud and Vertex AI surfaces according to the supplied research. Availability and preview status can vary by platform, so developers should check the relevant Google documentation before committing to a production integration.
Standard API pricing is listed as follows:
- Text and image input: $0.50 per 1 million tokens
- Batch text and image input: $0.25 per 1 million tokens
- Text and thinking output: $3 per 1 million tokens
- Image output: $60 per 1 million image-output tokens
- Batch image output: $30 per 1 million image-output tokens
Google’s equivalent image prices are approximately $0.045 for 512-pixel output, $0.067 for 1K, $0.101 for 2K, and $0.151 for 4K output. These image-equivalent figures are useful for estimating the cost of a generated asset, while token-based billing remains the underlying pricing method. Search grounding may incur separate charges after any applicable free allowance.
The cost and speed trade-off is central to the model’s positioning. Nano Banana 2 is intended for rapid iteration and high-volume generation, while larger or more specialized image systems may be preferable when the highest possible visual refinement matters more than throughput. The supplied editorial scores rate its speed at 9 out of 10 and cost at 7 out of 10; those scores are evaluations rather than provider benchmarks.
Main strengths and limitations
Strengths
- Fast image generation and editing for interactive workflows
- Text, image, video, and PDF understanding in one visual model
- Conversational multi-turn refinement
- 512px through 4K image output
- Very wide and very tall aspect-ratio options
- Support for Google Search and Image Search grounding
- Improved text rendering and multilingual visual content
- Provider-documented consistency for multiple characters and objects
Limitations
- No audio input, audio generation, video generation, or speech output
- No function calling, structured outputs, code execution, file search, URL context, or Google Maps grounding
- No Live API support or caching
- Exact image counts may not always be followed
- Text inside images may still need careful prompting and manual review
- Image Search grounding cannot currently use real-world images of people from web search
- Generated images include a SynthID watermark
- Search grounding can add charges beyond the applicable free allowance
When to choose Nano Banana 2
Choose Nano Banana 2 when the main deliverable is an image and the workflow benefits from fast iteration. It is a strong fit for marketing teams producing multiple asset variations, developers building image-editing features, designers creating posters or mockups, and applications that need to combine visual references with current search-grounded information.
It is especially suitable when a request involves more than a simple text-to-image prompt. A workflow that starts with a PDF brief, adds reference images, asks for a particular aspect ratio, and then performs several conversational edits matches the model’s intended role. The 512px-to-4K range also allows teams to choose a lower-cost output for previews and a higher-resolution version for final delivery.
Another model type may be more appropriate when the primary task is long-form reasoning, software development, strict JSON generation, external function execution, audio processing, or video creation. Even within an image pipeline, a higher-end image model may be preferable when maximum visual quality or specialized control is more important than Flash-level speed and cost efficiency. Nano Banana 2 is best understood as a fast visual production model, not a universal replacement for language, agent, audio, or video systems.
Bottom line
Nano Banana 2 combines image generation, image editing, multimodal input, search grounding, and broad output formats in a model designed for speed and repeated use. Its stable Gemini API identity is Gemini 3.1 Flash Image, and its practical advantages are most visible in visual workflows that require quick revisions, reference-aware generation, readable text, and output sizes up to 4K.
Its boundaries are equally important: it does not provide general tool use, structured JSON, code execution, audio, or video generation. For teams that need a responsive image model rather than a general-purpose agent, those trade-offs are reasonable. For applications centered on automation or text reasoning, Nano Banana 2 is more useful as a visual specialist paired with another model than as the sole system.

