Google DeepMind
Nano Banana 2.1
Image generation, reference pictures and conversational edits

Key facts
- 1K to 4K
- Output sizes
- Up to 14
- Reference images
- $0.0336Standard API, extra charges apply
- 1K image output
- Gemini and API
- Model access
Nano Banana 2.1 creates pictures from instructions and edits images you supply. Google's model can combine reference pictures, keep characters across successive edits and use web or image search for grounding.
Listen to the model guide
Podcast: Google's latest image editing model
Nano Banana 2.1 is Google’s AI model for making pictures from instructions and editing images you provide. Its 6 October 2026 model card describes text and image generation, with access through products including Gemini, AI Studio and Flow.
An editing conversation can start with an uploaded picture, then ask for a different background, a changed pose or a new composition. Reference pictures supply visual details the model can use across the result.
What has Google improved?
Google’s API documentation describes improvements in visual quality, following instructions, text rendering and character consistency across successive edits. Nano Banana 2.1 follows Nano Banana 2 in Google’s image-model family.
The model generates 1K, 2K and 4K images, with 1K as the default, settings also documented in Fal’s editing API schema. Google also reports improvements to wide and panoramic outputs at the higher resolutions, including the supported 1:4, 4:1, 1:8 and 8:1 aspect ratios.
These are the publisher’s stated changes. For a creative workflow, compare the outputs that affect the finished work: legible lettering, the position of objects, a subject’s appearance and the details carried over from the supplied images.
How many reference pictures can you combine?
Nano Banana 2.1 supports up to 14 reference images in its documented multi-image workflow. Google lists character consistency for up to four characters and object fidelity for up to ten objects.
A reference can establish a person’s appearance, a product’s shape or the style of a scene. An instruction then explains how those ingredients should fit together: for example, placing a photographed product into a different setting while retaining its visible design.
Google’s model card includes evaluations of multi-reference editing and character consistency. Its reported improvements describe those tests. A production review should compare the result with the supplied references, including proportions, logos, clothing and other identifying details that need to carry through.

What does an image cost through the API?
Google’s pricing page lists the following image-output equivalents on 6 October 2026. Prices are in US dollars and cover the generated image output.
| Output resolution | Standard API | Batch API |
|---|---|---|
| 1K | $0.0336 | $0.0168 |
| 2K | $0.0504 | $0.0252 |
| 4K | $0.0756 | $0.0378 |
Standard input costs $1.50 per million tokens; generated text and thinking cost $7.50 per million tokens. Those charges are added to the image-output charge. Search grounding has its own allowance and billing, and product subscriptions apply their own usage rules.
The distinction helps when estimating a batch of edits: count the image outputs, the supplied content and any generated text or thinking. Google’s published image equivalents describe one part of the final request bill.

Which controls can developers use?
The Gemini API identifies this model as gemini-nano-banana-2.1. Its endpoint documentation lists 131,072 input tokens and 32,768 output tokens, plus batch access and grounding with web and image search.
Developers can choose minimal, medium or high thinking, with medium as the default. The setting controls the model’s reasoning work around generation and editing. Search grounding can bring retrieved information or imagery into the request.
Google’s model card describes areas that need careful review, including small lettering, character consistency and left/right positioning. For an image intended for publication, check its text, subject details and composition at the size readers will see. Successive edits can then target the details that need adjustment.
More in Image Generation
All Image →- GoogleNano Banana Pro and Nano Banana 2the 4K value leader
- OpenAIGPT Image 2.5two variants, one price, the top two seats
- Black Forest LabsFLUX 3 ImageFLUX 3 Image creates and edits pictures with up to ten reference images
- OpenAIGPT Image 2the leading line
- AlibabaQwen-Image-2.1one small model for making pictures and editing them
- Black Forest LabsFLUX.2the strongest open-weight family