Gemini Image

Nano Banana 2

by Google DeepMind · Generally available; the stable Gemini API model is gemini-3.1-flash-image

Nano Banana 2 is Google DeepMind’s Gemini 3.1 Flash Image model for fast, high-volume image generation and editing. It accepts text, images, video, and PDFs, supports Google Search and Image Search grounding, offers 512px through 4K output, and includes configurable thinking. Its main limitations are the lack of function calling, structured outputs, code execution, audio, and video generation.

Text Image generation Reasoning Coding
Nano Banana 2 is the marketing name for Google DeepMind’s Gemini 3.1 Flash Image model. It is a native multimodal image model for creating and editing visuals through text conversation, with support for reference images, video, PDF inputs, image-search grounding, accurate text rendering, broad aspect ratios, and output from 512 pixels to 4K. Its main appeal is the combination of Flash-level speed and production-oriented image quality rather than general-purpose coding, tool use, or structured data generation.
Outputs

What Nano Banana 2 can produce

Text Image generation
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Web search Batch API Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
2/10 Coding
9/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Gemini Image
Model type Multimodal
Context window 131K tokens
Maximum output 33K tokens
Knowledge cutoff January 2025
Release date February 26, 2026
Status Generally available; the stable Gemini API model is gemini-3.1-flash-image
Knowledge cutoff notes

Google’s Gemini 3 developer documentation lists January 2025 as the knowledge cutoff for the Gemini 3.1 Flash Image model line. Search and image grounding can provide newer external information during use but do not change the underlying cutoff.

Model notes

Nano Banana 2 is the marketing name for Gemini 3.1 Flash Image. The stable API identifier is gemini-3.1-flash-image. Google introduced the model on February 26, 2026 and documented the generally available version on May 28, 2026. It supports 512px, 1K, 2K, and 4K image output, configurable thinking, Google Search and Image Search grounding, and broad aspect-ratio support. Google documents support for text, image, video, and PDF input and image/text output. Function calling, structured outputs, code execution, file search, URL context, Google Maps grounding, audio generation, Live API, and caching are listed as unsupported. Generated images include a SynthID watermark. Image Search grounding does not currently support using real-world images of people from web search.

Cost

Model pricing

Input $0.50 per 1 million tokens for text/image input; batch input $0.25 per 1 million tokens
Output $3 per 1 million text/thinking tokens and $60 per 1 million image-output tokens; approximately $0.067 per 1K image, $0.101 per 2K image, and $0.151 per 4K image; batch image output $30 per 1 million image-output tokens
Model guide

Nano Banana 2: Fast Gemini Image Generation and Editing at up to 4K

Nano Banana 2 is Google DeepMind’s Gemini 3.1 Flash Image model, built for fast, high-volume image generation and conversational editing. It accepts text, images, video, and PDFs, supports image-search grounding and 512px-to-4K output, and trades some general-purpose language and tool capabilities for lower latency and efficient visual production.

What is Nano Banana 2?

Nano Banana 2 is Google DeepMind’s marketing name for Gemini 3.1 Flash Image, a native multimodal model designed primarily for image generation and image editing. Its stable Gemini API model identifier is gemini-3.1-flash-image.

Unlike a text-only language model, Nano Banana 2 can understand visual material and produce images as part of a conversational workflow. A user can provide a written instruction, one or more reference images, video, or a PDF, then ask the model to create, modify, translate, localize, upscale, or recompose visual content. Follow-up prompts can refine the result without restarting the entire task.

Within Google’s Nano Banana family, this model is positioned as the general-purpose, high-throughput option. It aims to retain many of the control and quality features associated with higher-end image generation while offering the speed expected from a Flash model. That positioning makes it more suitable for repeated visual iteration and production pipelines than for open-ended text reasoning or agentic automation.

What Nano Banana 2 can create and edit

The model supports both text-to-image generation and conversational editing. Practical requests can include creating a marketing poster, changing the style of a product photograph, translating text within an image, turning a video frame into a cinematic thumbnail, or combining information from a PDF with a visual design brief.

  • New images generated from text prompts
  • Edits based on one or more reference images
  • Multi-turn refinement of an existing image
  • Posters, social media graphics, mockups, diagrams, and infographics
  • Image localization and translation
  • Visual summaries based on video or PDF material
  • Current or location-specific visuals using Google Search and Image Search grounding

Google documents support for 512-pixel, 1K, 2K, and 4K image outputs. The model also supports unusually wide and tall formats, including 1:4, 4:1, 1:8, and 8:1, alongside more conventional aspect ratios. These formats are useful for banners, mobile layouts, panoramic graphics, and other designs that do not fit a standard landscape or portrait frame.

Inputs, outputs, and visual understanding

Nano Banana 2 accepts text, images, video, and PDF files. Video can provide visual context for tasks such as selecting a scene for a thumbnail, designing a cinematic poster, or producing a summary infographic. Audio input is not supported for this image-generation model.

The model can return both generated images and text. The image is the important non-text output: accompanying text may describe the result or provide a response to the instruction, but Nano Banana 2 is not primarily intended as a conversational text model. It does not generate video, audio, music, speech, or embeddings.

Google also highlights improved multilingual text rendering. This is relevant when an image needs labels, headlines, signs, product copy, or other readable language. Text inside generated images can still require careful prompting and verification. Google notes that results may be better when the desired wording is specified precisely or developed before the rendering step rather than treated as an afterthought.

Grounding, consistency, and creative control

Nano Banana 2 can use Google Search and Google Image Search grounding. Grounding means that the model can consult current external information or visual references during a request, which is useful for subjects such as recent products, current events, locations, and factual visual explanations. Grounding is an API or platform feature; it does not change the model’s underlying knowledge cutoff, which is listed as January 2025 for this model line.

The model supports configurable thinking behavior, improved instruction following, stronger aspect-ratio adherence, and visual consistency across an iterative workflow. Google reports that it can preserve the resemblance of up to four characters and the fidelity of up to ten objects in a single workflow. Those figures are provider claims rather than an independent guarantee for every prompt or composition.

Generated images include a SynthID watermark. Google also documents a limitation on Image Search grounding: it does not currently support using real-world images of people from web search. As with other generative image systems, users should review outputs for factual accuracy, unwanted alterations, illegible text, and resemblance issues before publication.

Technical specifications at a glance

SpecificationDocumented detail
Model nameGemini 3.1 Flash Image
Marketing nameNano Banana 2
Stable model IDgemini-3.1-flash-image
Context length131,072 tokens
Maximum output tokens32,768
Input typesText, image, video, and PDF
Output typesImage and text
Image resolutions512px, 1K, 2K, and 4K
Image groundingGoogle Search and Google Image Search
Fine-tuningNot supported according to the supplied model documentation
CachingNot supported according to the supplied model documentation

The context length describes the total amount of model context that can be supplied to a request, including relevant text and multimodal material as measured by the platform. It should not be interpreted as a promise that every large collection of files will produce equally strong results. The maximum output-token value is a platform specification for generated text and thinking output; image dimensions and image-output token accounting are separate concerns.

Reasoning, coding, and tool support

Nano Banana 2 includes configurable thinking behavior, which can help with complex visual instructions, composition decisions, and multi-step editing. However, it is still an image-focused model rather than a general reasoning model. The supplied editorial evaluation rates its reasoning at 7 out of 10 and coding at 2 out of 10; these are subjective catalog scores, not Google-published benchmark results.

The model documentation lists function calling, structured outputs, code execution, file search, URL context, Google Maps grounding, audio generation, and Live API support as unavailable. This distinction matters when selecting an API model. Nano Banana 2 can use Search or Image Search grounding for supported visual generation tasks, but it is not intended to operate as a tool-calling agent that plans and executes business actions.

For example, it may be appropriate to create a location-themed poster using grounded information, but not to build an assistant that reliably calls a booking system, returns strict JSON, executes code, or manages a multi-step workflow through external functions. A separate language or agent model is more appropriate for those jobs, with Nano Banana 2 used as the visual generation component when needed.

API pricing and availability

Nano Banana 2 is available through the Gemini API and Google AI Studio, with enterprise deployment through Google Cloud and Vertex AI surfaces according to the supplied research. Availability and preview status can vary by platform, so developers should check the relevant Google documentation before committing to a production integration.

Standard API pricing is listed as follows:

  • Text and image input: $0.50 per 1 million tokens
  • Batch text and image input: $0.25 per 1 million tokens
  • Text and thinking output: $3 per 1 million tokens
  • Image output: $60 per 1 million image-output tokens
  • Batch image output: $30 per 1 million image-output tokens

Google’s equivalent image prices are approximately $0.045 for 512-pixel output, $0.067 for 1K, $0.101 for 2K, and $0.151 for 4K output. These image-equivalent figures are useful for estimating the cost of a generated asset, while token-based billing remains the underlying pricing method. Search grounding may incur separate charges after any applicable free allowance.

The cost and speed trade-off is central to the model’s positioning. Nano Banana 2 is intended for rapid iteration and high-volume generation, while larger or more specialized image systems may be preferable when the highest possible visual refinement matters more than throughput. The supplied editorial scores rate its speed at 9 out of 10 and cost at 7 out of 10; those scores are evaluations rather than provider benchmarks.

Main strengths and limitations

Strengths

  • Fast image generation and editing for interactive workflows
  • Text, image, video, and PDF understanding in one visual model
  • Conversational multi-turn refinement
  • 512px through 4K image output
  • Very wide and very tall aspect-ratio options
  • Support for Google Search and Image Search grounding
  • Improved text rendering and multilingual visual content
  • Provider-documented consistency for multiple characters and objects

Limitations

  • No audio input, audio generation, video generation, or speech output
  • No function calling, structured outputs, code execution, file search, URL context, or Google Maps grounding
  • No Live API support or caching
  • Exact image counts may not always be followed
  • Text inside images may still need careful prompting and manual review
  • Image Search grounding cannot currently use real-world images of people from web search
  • Generated images include a SynthID watermark
  • Search grounding can add charges beyond the applicable free allowance

When to choose Nano Banana 2

Choose Nano Banana 2 when the main deliverable is an image and the workflow benefits from fast iteration. It is a strong fit for marketing teams producing multiple asset variations, developers building image-editing features, designers creating posters or mockups, and applications that need to combine visual references with current search-grounded information.

It is especially suitable when a request involves more than a simple text-to-image prompt. A workflow that starts with a PDF brief, adds reference images, asks for a particular aspect ratio, and then performs several conversational edits matches the model’s intended role. The 512px-to-4K range also allows teams to choose a lower-cost output for previews and a higher-resolution version for final delivery.

Another model type may be more appropriate when the primary task is long-form reasoning, software development, strict JSON generation, external function execution, audio processing, or video creation. Even within an image pipeline, a higher-end image model may be preferable when maximum visual quality or specialized control is more important than Flash-level speed and cost efficiency. Nano Banana 2 is best understood as a fast visual production model, not a universal replacement for language, agent, audio, or video systems.

Bottom line

Nano Banana 2 combines image generation, image editing, multimodal input, search grounding, and broad output formats in a model designed for speed and repeated use. Its stable Gemini API identity is Gemini 3.1 Flash Image, and its practical advantages are most visible in visual workflows that require quick revisions, reference-aware generation, readable text, and output sizes up to 4K.

Its boundaries are equally important: it does not provide general tool use, structured JSON, code execution, audio, or video generation. For teams that need a responsive image model rather than a general-purpose agent, those trade-offs are reasonable. For applications centered on automation or text reasoning, Nano Banana 2 is more useful as a visual specialist paired with another model than as the sole system.


Answers to Frequently Asked Questions

What are the main limitations of Nano Banana 2?
Nano Banana 2 does not support function calling, structured outputs, code execution, file search, URL context, Google Maps grounding, Live API, caching, audio, speech, or video generation. Generated images include a SynthID watermark, and image text, exact image counts, factual accuracy, and resemblance should be reviewed before publication.
How much does Nano Banana 2 cost through the API?
Standard pricing includes $0.50 per 1 million text and image input tokens, $3 per 1 million text and thinking output tokens, and $60 per 1 million image-output tokens. Approximate image-equivalent prices are $0.045 for 512-pixel output, $0.067 for 1K, $0.101 for 2K, and $0.151 for 4K output. Search grounding may add separate charges.
Can Nano Banana 2 generate 4K images and wide aspect ratios?
Yes. The model supports 512-pixel, 1K, 2K, and 4K image outputs. It also supports very wide and tall formats, including 1:4, 4:1, 1:8, and 8:1 aspect ratios.
What is Nano Banana 2?
Nano Banana 2 is Google DeepMind’s marketing name for Gemini 3.1 Flash Image, a multimodal model focused on image generation and editing. Its stable Gemini API model identifier is gemini-3.1-flash-image.
What types of inputs and outputs does Nano Banana 2 support?
Nano Banana 2 accepts text, images, video, and PDF files. It can return generated images and text, but it does not generate video, audio, music, speech, or embeddings.


Sources 8
Provider

About Google DeepMind