What is Nano Banana Pro?
Nano Banana Pro is the consumer-facing name for Google DeepMind’s Gemini 3 Pro Image model. It is a multimodal model: users can provide text and images, and the model can return both text and generated images. This supports conversational creation and editing rather than limiting the model to one-shot text-to-image prompts.
The model is positioned as the professional-focused member of Google’s current Nano Banana image-generation family. Its intended role is to handle visual tasks that require more than a simple illustration, including complex compositions, product visualization, branded assets, diagrams, factual infographics, image localization, and iterative editing based on several references.
Google’s model documentation identifies gemini-3-pro-image as the canonical API model ID. Product access can differ depending on whether the model is used through the Gemini API, Google AI Studio, or a Google consumer or business product. Availability, quotas, pricing, and feature access may therefore depend on the specific Google service and account tier.
What can Nano Banana Pro do?
Nano Banana Pro combines image generation with image understanding and editing. A user can describe a new scene, provide an existing image for modification, or submit multiple reference images and ask the model to combine their important characteristics.
- Generate images from written descriptions.
- Edit existing images using natural-language instructions.
- Create product mockups and branded marketing assets.
- Produce diagrams, charts, infographics, and educational explainers.
- Render readable or stylized text directly inside images.
- Localize visual material into different languages or markets.
- Combine multiple object, character, and style references.
- Change camera angle, framing, lighting, color, depth of field, or scene scale.
- Generate images at 1K, 2K, and 4K resolutions.
These capabilities make the model more suitable for controlled visual production than for quick, disposable image variations. For example, a product team could provide a product reference, a packaging reference, and a campaign style reference, then request a specific advertising composition while preserving the product’s recognizable appearance.
Reasoning and Google Search grounding
Nano Banana Pro includes a thinking process intended to improve complex visual compositions before the final image is produced. In practical terms, this is relevant when a prompt specifies several objects, relationships, layout requirements, visual styles, or factual details that must work together.
The model also supports Google Search grounding when that feature is enabled. Search grounding supplies information from current web results during a request, which can help with images involving changing facts such as recent events, sports, weather, market information, recipes, or educational subjects. It does not change the model’s underlying knowledge cutoff; it provides additional information for the particular generation request.
Search grounding should not be confused with general-purpose autonomous tool use. The supplied model documentation does not list function calling, code execution, file search, URL context, or Live API support for Nano Banana Pro.
Reference images and editing control
Nano Banana Pro supports image-to-image workflows and multi-reference composition. Google’s image-generation documentation lists support for up to 14 total reference images for Gemini 3 image models. For the documented high-fidelity configuration, Nano Banana Pro supports up to six object references, five character references, and three style references.
Reference images can be used to preserve identity and visual consistency across related assets. Appropriate examples include maintaining the appearance of a product across advertising scenes, keeping a character consistent across storyboard frames, applying a campaign style to new layouts, or adapting an advertisement for another language while retaining the original design direction.
Natural-language editing can also target specific visual properties. Prompts may request a different framing, aspect ratio, lighting setup, color treatment, focus, camera angle, or scene context. Results still need to be checked, especially when an image contains dense text, strict object counts, or a layout that must match production specifications exactly.
Nano Banana Pro technical specifications
| Specification | Documented value |
|---|---|
| Canonical API model ID | gemini-3-pro-image |
| Input modalities | Text and image |
| Output modalities | Text and image |
| Input token limit | 65,536 tokens |
| Maximum output token limit | 32,768 tokens |
| Image resolutions | 1K, 2K, and 4K |
| Google Search grounding | Supported |
| Thinking | Supported |
| Batch API | Supported |
| Function calling | Not supported |
| Structured outputs | Not supported |
| Code execution | Not documented for this model |
| Context caching | Not supported |
| Live API | Not supported |
The token limits describe the model’s text-processing capacity, not the number of pixels in an image. The image output choices are separately expressed as 1K, 2K, or 4K resolution tiers. Nano Banana Pro can return text as well as an image, but it is not an audio or video generation model.
Pricing and access
Google’s standard paid-tier Gemini API pricing lists text and image input at $2.00 per 1 million tokens. Google’s documentation also gives an approximate input cost of $0.0011 per input image. Text and thinking output tokens are priced at $12.00 per 1 million tokens.
Image output is listed at $120.00 per 1 million image tokens. The documented approximate per-image equivalents are $0.134 for a 1K or 2K image and $0.24 for a 4K image. Batch and Flex processing offer lower rates, with image output listed at approximately $0.067 for a 1K or 2K image and $0.12 for a 4K image. Google Search grounding can add separate charges after the applicable monthly free allowance.
These are API pricing figures rather than a universal subscription price for every Google product. A Gemini application may package model access into a plan, impose usage limits, or expose a different set of controls. Developers should check the pricing and quota rules for the specific endpoint and processing mode they intend to use.
Main strengths and trade-offs
The strongest case for Nano Banana Pro is high-fidelity visual work that benefits from detailed instructions and several constraints. Its combination of reasoning, reference-image support, text rendering, high-resolution output, and image editing makes it a better fit for professional creative tasks than a model selected solely for fast, inexpensive variations.
- Complex composition: The model is designed to coordinate multiple visual elements and relationships in one request.
- Text in images: It is intended for designs that require legible labels, headlines, diagrams, charts, or other embedded text.
- Reference control: Multiple object, character, and style references can help preserve a visual direction across assets.
- Resolution: 1K, 2K, and 4K output options support different quality and delivery requirements.
- Current information: Optional Search grounding can support visualizations that depend on changing factual information.
- Professional editing: Follow-up instructions can adjust an existing composition instead of requiring a complete restart.
The principal trade-off is cost and likely throughput. Nano Banana Pro is more expensive than Google’s Flash-oriented image alternatives, so it may not be the most economical choice for bulk generation, rapid experimentation, or applications where a small difference in composition quality is acceptable. Generated images also include Google’s SynthID watermarking technology.
Coding, tools, and automation support
Nano Banana Pro is primarily a visual model, not a general-purpose coding or agent model. Its documented API capabilities do not include function calling, code execution, file search, URL context, Live API access, structured outputs, or context caching for this exact model.
That limitation matters when integrating it into an automated workflow. The model can generate visual assets and return text, but an application should not assume that it can call business tools, execute code, retrieve arbitrary files, or produce schema-validated structured responses. Developers needing those behaviors may need to place the image model inside a separate application orchestration layer or select a different model for the tool-using part of the workflow.
The supplied evaluation data gives Nano Banana Pro a coding score of 2, reasoning score of 8, speed score of 6, and cost score of 5. These are editorial or database assessments, not provider-published benchmark results. They indicate that the model is best understood as a reasoning-oriented image model rather than a coding specialist, while its cost and speed occupy a middle position relative to the broader model catalog.
Best use cases
Nano Banana Pro is a strong candidate when visual quality, controllability, and fidelity are more important than the lowest unit price. Suitable applications include:
- Professional marketing and advertising imagery.
- High-fidelity product visualization and packaging mockups.
- Brand-consistent campaign asset production.
- Infographics, diagrams, and educational visual explainers.
- Multilingual creative localization.
- Concept development and visual prototyping.
- Storyboards and reference-driven scene design.
- Complex editing involving several source images.
It is less appropriate for audio or video generation, general-purpose software development, function-calling agents, structured-data workflows, or extremely cost-sensitive image generation at very high volume.
When should you choose Nano Banana Pro?
Choose Nano Banana Pro when a request requires a carefully composed image, accurate visual text, multiple references, high-resolution delivery, or iterative editing with detailed instructions. It is particularly suitable when a creative team is willing to pay more per output to reduce manual correction and improve consistency across a set of professional assets.
Choose a faster or cheaper image model when the task is simple, exploratory, or high volume and does not require 4K output, extensive reference control, or complex layout reasoning. Within Google’s Nano Banana family, Nano Banana 2 is positioned as the more versatile Flash-oriented option, while Nano Banana 2 Lite targets lower latency and cost. Those alternatives may be more appropriate for rapid previews, routine variations, or workloads where throughput matters more than maximum fidelity.
Nano Banana Pro should also be paired with another model or application layer when the workflow requires tool calls, code execution, structured JSON, file search, or real-time interaction. Its value is concentrated in high-quality image creation and editing, not in serving as a complete autonomous agent.
Limitations to check before production
Before deploying Nano Banana Pro in a production pipeline, validate the exact image sizes, quotas, pricing mode, and access conditions for the Google product being used. API availability does not automatically guarantee identical access in consumer applications.
Image-generation behavior can vary with prompt structure, reference quality, aspect ratio, and the complexity of requested text or composition. Even with strong text rendering and reference support, exact object counts and intricate layouts should be reviewed rather than assumed to be pixel-perfect. Teams should also account for SynthID watermarking and the model’s lack of documented audio, video, function-calling, structured-output, and caching features.
Overall, Nano Banana Pro is best viewed as Google’s premium image-production option for complex, reference-driven visual work. Its higher cost is justified when composition quality, resolution, embedded text, and editing control are central requirements; a Flash-oriented sibling is likely a better choice when speed and cost dominate.

