On October 6, Google made its image generation and editing model Nano Banana 2.1 generally available. Compared with the previous Nano Banana 2, the API price for image output has roughly halved. In Google's internal evaluations, the model also outperforms Nano Banana Pro on several measures, including character consistency and partial editing. The update aims to reduce a familiar problem: you want to fix only the background or the text in an ad, but a person's face or a product's shape changes as well.

However, what got cheaper is image output; unit prices for input and reasoning went up. When comparing with models from OpenAI or Black Forest Labs, you need to look beyond the per-image price to editing controllability and the image sizes each model can output.

Nano Banana 2.1 is built on Gemini 3.6 Flash, and its official API name is gemini-nano-banana-2.1. Google says it has improved image quality and instruction following, and strengthened text rendering, diagram layout, and character consistency across repeated edits.

For example, when a product ad is translated into another language and the background is swapped, how well the product's appearance and the people in it are preserved is what matters. Google also fixed an issue in which the same pattern repeated like tiles in extremely wide or tall 2K and 4K images. Extreme aspect ratios were already available in the old model, so this is an improvement in output quality rather than the addition of a new supported format.

AD

Image output about half the price, but input and reasoning rates rise

Standard API pricing for Nano Banana 2.1 is $0.0336 per 1K image, $0.0504 per 2K image, and $0.0756 per 4K image. Comparing Google's image models in the same pricing tiers, the 1K image output price of 2.1 now matches that of Nano Banana 2 Lite.

Google model 1K image output 2K image output 4K image output
Nano Banana 2.1 $0.0336 $0.0504 $0.0756
Nano Banana 2 $0.067 $0.101 $0.151
Nano Banana 2 Lite $0.0336 Not supported Not supported
Nano Banana Pro $0.134 $0.134 $0.24
Original Nano Banana $0.039 Not supported Not supported

Source: Gemini API pricing. Standard rates as of October 7, 2026, comparing image output only, in US dollars. Prices for older models and Pro use the rounded per-unit figures listed in the official pricing table. Input, reasoning, search and other charges are not included.

Generating 10,000 1K images through the standard API would cost $336 in image output fees, down from $670 with the old model, a saving of $334. This is an approximation based on the figures shown in the official pricing table: 0.067 × 10,000 = $670 and 0.0336 × 10,000 = $336. Calculated strictly from token prices, the old model would come to $672, but either way image output costs fall by nearly half.

For the Batch API, designed for bulk processing, prices are halved again: $0.0168 per 1K image and $0.0378 per 4K image with 2.1.

That said, your total API bill will not necessarily fall by the same proportion. Input pricing per million tokens rose from $0.50 for the old Nano Banana 2 to $1.50 for 2.1. Output pricing for text and reasoning also rose from $3 to $7.50. In workflows that send reference images and revise repeatedly, these processing costs add up on top of the image output fees themselves.

Nano Banana 2 Lite still has a role. Its 1K image output price is the same as 2.1's, but Lite's input price is far lower at $0.25 per million tokens, and text/reasoning output is $1.50.

Google positions Lite as a model focused on speed and high-volume processing; it is not optimized for complex production involving multiple reference images or repeated editing rounds. It also does not support search grounding. Even with identical image output prices, there are different reasons to choose each model: light-use cases where you generate once and finish, versus work where you refine repeatedly while keeping people and products intact.

How to read Google's internal evaluations showing it beats Pro

In the evaluations Google published in its October model card, Nano Banana 2.1 outperformed both the previous Nano Banana 2 and Nano Banana Pro in both image generation and editing. The table below pulls out the metrics most likely to make a difference in real production work.

Google evaluation metric 2.1, with reasoning 2.1, without reasoning Old 2, with reasoning Pro
Overall preference 1050±14 1015±13 990±7 935±8
Diagram design 1048±17 1001±17 961±12 912±12
Diagram factuality 0.521 0.328 0.179 0.265
General image editing 1026±12 980±15 938±11 939±10
Multi-character consistency 1106±14 1068±14 978±10 1011±10
Editing via masks / hand-drawn instructions 1049±15 1042±16 965±12 927±12

Source: Google DeepMind model card. Evaluations published by Google in October 2026. "Diagram factuality" is an automated evaluation; the rest are Elo scores based on human side-by-side comparisons of two images. The ± notation from the original table is retained.

The large gains in character consistency and partial-editing scores match the improvements Google describes this time. In tasks such as changing only one person's clothing in a group photo or swapping just the background while leaving the product untouched, how accurately the model separates "what to change" from "what to keep" determines how many retries are needed.

However, this table does not include GPT Image 2.5 or FLUX 3. Comparisons among Google's own models cannot establish who leads the market once other companies' models are included.

Elo score differences also cannot be translated into statements like "X% higher image quality" or "N times faster generation." Google does not describe the 0.521 for "diagram factuality" as a simple 52.1% accuracy rate either. Producing diagrams that people find appealing is a different question from whether the information depicted in them is factually correct.

Caution is also needed on reasoning. The model card has a "without reasoning" column, but the public API does not offer a setting that disables reasoning entirely. Nano Banana 2.1 lets you choose among three levels, minimal, medium and high, with medium as the default. The "without reasoning" condition in the evaluation environment should not be taken as an API setting users can select.

Nano Banana Pro does not become unnecessary overnight. Google still positions Pro for uses that require advanced knowledge, multilingual support, brand consistency and precise production control.

Reference image handling differs as well. Google says 2.1 can reference up to 10 objects and 4 people, while Pro can reference 6 objects, 5 people and 3 styles. Simple totals of image counts do not allow a comparison of how faithfully people or products are reproduced.

Google itself acknowledges that 2.1 still has limitations. In 1K images, small text and long passages tend to break down, and matching people is not perfect. Lines used for hand-drawn editing instructions can remain in the finished image, and left-right positional relationships can be confused. For uses that put fine text directly into images, such as Japanese product descriptions, you still need a step to proofread the actual text after generation.

AD

Versus GPT Image 2.5: transparent backgrounds and what "4K" means

OpenAI's current image models include "GPT Image 2.5 Sunburst," which emphasizes editing precision, and "GPT Image 2.5 Flare," which emphasizes speed. OpenAI says both support image generation and editing and officially support transparent backgrounds.

Because backgrounds can be made transparent in PNG or WebP, this is a concrete point of difference from Google's models for uses such as cut-out product assets or components to be layered into other designs later.

Pricing has to be viewed differently from Google's. Image output for both Sunburst and Flare is $30 per million tokens under standard processing. Nano Banana 2.1's image output is also $30 per million tokens, but the number of tokens consumed to generate one image differs by model.

With OpenAI you can choose quality settings such as low, medium, high, xhigh and max, and token consumption changes with the setting. Even at the same "$30 per million tokens," the per-image price is not necessarily the same.

The label "4K" cannot be compared directly either. A 4K square image on Google is 4,096 × 4,096 pixels, for a total of 16,777,216 pixels.

GPT Image 2.5, by contrast, caps total pixels at 8,294,400, with a maximum edge of 3,840 pixels and aspect ratios from 1:3 to 3:1.

In other words, although both are called "4K," Nano Banana 2.1's 4K square is about 16.78 million pixels while GPT Image 2.5's maximum is about 8.29 million pixels, so actual image sizes differ. The former is a 4K square specification; the latter is an upper limit on total pixels applied across various aspect ratios, and it does not mean image quality is twice as high.

The figures are total pixel counts calculated from Google's dimension table and OpenAI's size constraints as of October 7, 2026. OpenAI treats resolutions above 2,560 × 1,440 as experimental.

For wide signage and banners, the difference grows. Nano Banana 2.1 can output 12,288 × 1,536 pixels at an 8:1 4K setting, while OpenAI limits aspect ratios to 3:1 at most.

So rather than asking only "does it support 4K," choosing a model based on the final aspect ratio and pixel dimensions you need makes it easier to judge the cropping and upscaling work required after generation.

FLUX 3 differentiates on region specification, Qwen on self-hosting

Black Forest Labs' "FLUX 3 Image" can combine up to 10 reference images and can also use information and images retrieved from web search in generation. One of its features is that you can specify rectangular regions to determine where objects are placed. The same mechanism can be used to edit the color or shape of just one part of an image.

When you want to move one part while preserving a person's face or a product's appearance, the difference in controllability is whether you instruct with text alone or can also specify location explicitly.

Lining up the features and terms of the image models shows that competition cannot be compared by simple image quality scores alone.

Model Resolution / reference characteristics Features and conditions to compare for production
Nano Banana 2.1 1K, 2K, 4K; up to 14 reference images Web and image search, consistency in repeated editing, extremely wide or tall images
Nano Banana Pro 1K, 2K, 4K; references including 5 people and 3 styles Emphasis on advanced knowledge, multilingual support, brand consistency and precise production control
Nano Banana 2 Lite Up to 1K; 2K and 4K not supported Low input and reasoning rates, for speed and bulk processing, no search support
GPT Image 2.5 Sunburst / Flare Up to about 8.29 million pixels; aspect ratio up to 3:1 Choice of editing-precision or speed focus, transparent backgrounds, multiple quality levels
FLUX 3 Image 1K, 2K, 4K; up to 10 reference images Web and image search, placement and editing via rectangular region specification
Qwen-Image-2.1 Native 2K; up to 10 reference images Transparent image generation and editing, weights available for self-hosting, separate license required for commercial use

Sources: Google image generation guide, OpenAI image generation guide, FLUX 3 specifications, official Qwen implementation, Qwen weights license. This summarizes specifications and terms as of October 7, 2026; it is not a test that compared image quality using identical prompts. The number of referenceable images is also not a guarantee that every subject will be perfectly preserved.

Official FLUX 3 pricing is $0.048 per roughly 1-megapixel 1K image, $0.100 per roughly 4-megapixel 2K image, and $0.607 per roughly 16-megapixel 4K image.

On image output pricing alone, Nano Banana 2.1 is cheaper than these. However, because Google bills input, reasoning and other items separately, this gap cannot be treated as the overall savings rate for production. There are also lower prices elsewhere: FLUX.2 [klein] 4B, aimed at high-volume processing, starts at $0.014 per image, so calling Nano Banana 2.1 the cheapest model in the entire image generation market would not be accurate.

Even with models that let you specify the region to edit, the area outside the specified range is not necessarily fully fixed. BFL's editing guide says that while areas outside the specified region are normally preserved, shadows, reflections and surrounding lighting may change.

OpenAI likewise explains that repeated editing can alter details you want to keep, and for parts that must match exactly it recommends compositing the approved edit result back onto the original image. Google's high ratings for mask and hand-drawn editing also do not mean pixels outside the specified area are perfectly locked.

Qwen-Image-2.1 is an option of a different character altogether. Released on September 20, this model can directly generate and edit RGBA images with transparency, and also supports cutting out subjects from photos. It also supports editing with up to 10 reference images.

Because the model weights are public, it can be run on your own computing environment. However, the current Qwen Research License limits use to non-commercial research and evaluation, and a separate license must be obtained for commercial use. GPUs and operating environments also cost money, so being able to download the weights for free is not the same as being able to generate images for free.

AD

Settings to review for the October 29 migration

Google has announced that it will shut down the old Nano Banana 2 API name gemini-3.1-flash-image on October 29, and recommends migrating to Nano Banana 2.1. For Nano Banana Pro and Nano Banana 2 Lite, no shutdown date had been announced as of October 7. The entire Nano Banana series is not being replaced by 2.1 all at once.

If you are migrating from the old Nano Banana 2, you should also check that 0.5K output, at roughly 512 pixels, is not available in 2.1.

Reasoning settings change too. The old Nano Banana 2 defaulted to minimal reasoning, whereas 2.1 defaults to medium. If you change only the model name specified in the API and assume the same latency and processing cost as before, your actual estimates may be off.

To compare models in practice, it is easiest to use the materials you normally work with. Check whether Japanese headlines are rendered correctly and whether products and people are maintained when backgrounds or wording change. Repeating edits through background changes, text fixes and resizing will also reveal differences that a single generation will not.

What ultimately matters is not the unit price for generating one image but the cost and time to reach a finished image you can actually use. If revisions can be stacked while keeping people and products intact, and the number of retries drops, Nano Banana 2.1's price cut will have a greater effect in workplaces that produce ads and product images in volume.