YFarmX logoYFarmX

Google DeepMind

Nano Banana 2.1

Image generation, reference pictures and conversational edits

Released 6 October 20263 min readImage Generation

Editorial illustration of a photographic contact sheet and editing tools with Google branding, headed Nano Banana 2.1.

Key facts

1K to 4K
Output sizes
Up to 14
Reference images
$0.0336Standard API, extra charges apply
1K image output
Gemini and API
Model access

Nano Banana 2.1 creates pictures from instructions and edits images you supply. Google's model can combine reference pictures, keep characters across successive edits and use web or image search for grounding.

Listen to the model guide

Podcast: Google's latest image editing model

Nano Banana 2.1 is Google’s AI model for making pictures from instructions and editing images you provide. Its 6 October 2026 model card describes text and image generation, with access through products including Gemini, AI Studio and Flow.

An editing conversation can start with an uploaded picture, then ask for a different background, a changed pose or a new composition. Reference pictures supply visual details the model can use across the result.

What has Google improved?

Google’s API documentation describes improvements in visual quality, following instructions, text rendering and character consistency across successive edits. Nano Banana 2.1 follows Nano Banana 2 in Google’s image-model family.

The model generates 1K, 2K and 4K images, with 1K as the default, settings also documented in Fal’s editing API schema. Google also reports improvements to wide and panoramic outputs at the higher resolutions, including the supported 1:4, 4:1, 1:8 and 8:1 aspect ratios.

These are the publisher’s stated changes. For a creative workflow, compare the outputs that affect the finished work: legible lettering, the position of objects, a subject’s appearance and the details carried over from the supplied images.

How many reference pictures can you combine?

Nano Banana 2.1 supports up to 14 reference images in its documented multi-image workflow. Google lists character consistency for up to four characters and object fidelity for up to ten objects.

A reference can establish a person’s appearance, a product’s shape or the style of a scene. An instruction then explains how those ingredients should fit together: for example, placing a photographed product into a different setting while retaining its visible design.

Google’s model card includes evaluations of multi-reference editing and character consistency. Its reported improvements describe those tests. A production review should compare the result with the supplied references, including proportions, logos, clothing and other identifying details that need to carry through.

Reference-image diagram showing Nano Banana 2.1's documented allowance of up to 14 images, with character consistency for up to four characters and object fidelity for up to ten objects.
Google documents up to 14 reference images, with character consistency for up to four characters and object fidelity for up to ten objects.

What does an image cost through the API?

Google’s pricing page lists the following image-output equivalents on 6 October 2026. Prices are in US dollars and cover the generated image output.

Output resolution Standard API Batch API
1K $0.0336 $0.0168
2K $0.0504 $0.0252
4K $0.0756 $0.0378

Standard input costs $1.50 per million tokens; generated text and thinking cost $7.50 per million tokens. Those charges are added to the image-output charge. Search grounding has its own allowance and billing, and product subscriptions apply their own usage rules.

The distinction helps when estimating a batch of edits: count the image outputs, the supplied content and any generated text or thinking. Google’s published image equivalents describe one part of the final request bill.

Three photographic output cards showing standard API image-output equivalents of 0.0336 dollars at 1K, 0.0504 dollars at 2K and 0.0756 dollars at 4K.
Standard API image-output equivalents on 6 October 2026. Input, text, thinking and applicable search charges add to the request cost. Source: Google.

Which controls can developers use?

The Gemini API identifies this model as gemini-nano-banana-2.1. Its endpoint documentation lists 131,072 input tokens and 32,768 output tokens, plus batch access and grounding with web and image search.

Developers can choose minimal, medium or high thinking, with medium as the default. The setting controls the model’s reasoning work around generation and editing. Search grounding can bring retrieved information or imagery into the request.

Google’s model card describes areas that need careful review, including small lettering, character consistency and left/right positioning. For an image intended for publication, check its text, subject details and composition at the size readers will see. Successive edits can then target the details that need adjustment.