In image generation, Google has been quiet for a very long time. Even though the earlier Nano Banana 2 and Nano Banana Pro had pretty good generation results, they were still a notch behind ChatGPT, which led more and more people to jump ship to ChatGPT. Google finally seems ready to make a move: it suddenly released the Nano Banana 2.1 image generation model earlier, which is not only stronger than the previous-generation Nano Banana, but the official test data it shared also surpasses the higher-end Pro model, and the price is even lower—generating a 1K image via API costs only $0.0336.
It is now available in Gemini.

Google launches Nano Banana 2.1: image generation quality and character consistency improve further, API image generation price cut in half
Google announced the launch of Nano Banana 2.1 earlier through its official X account, and it also went live in the Gemini API the same day. This time, it skipped the preview phase, and the model code is gemini-nano-banana-2.1. In its post, Google said this upgraded version surpasses the previous model in every respect, with the most noticeable improvements in visual design, mask-based retouching, and subject consistency, and the generated images are also more natural.
Meet Nano Banana 2.1, our latest image generation and editing model 🍌 This upgraded version outperforms our previous models across the board, with notable leaps in visual design, mask-based editing, and subject consistency to help you create more natural-looking images.
— Google (@Google) October 6, 2026
According to the model card released by DeepMind, Nano Banana 2.1 belongs to the Gemini 3 series, and its underlying model is Gemini 3.6 Flash. It is officially positioned as the successor to Nano Banana 2, maintaining Flash-level speed and cost while also being more resource-efficient than Nano Banana Pro, so Pro has not been replaced.
There are four main points to this upgrade:
- Improved image quality and prompt adherence
- Role consistency is more stable in multi-turn conversations.
- More accurate text rendering and infographic layout
- Supports 1K, 2K, and 4K output, and mitigates stitching artifacts that can occur with ultra-wide and ultra-tall aspect ratios such as 1:4, 4:1, 1:8, and 8:1 at 2K and 4K.
For most of its generation and editing capabilities this time, Google used blind human side-by-side tests and then converted the results into Elo scores; the higher the score, the more often it was preferred by evaluators. Infographic accuracy, however, was evaluated separately using AutoRater, so those scores cannot be directly compared with Elo.
In the “Overall Preference” category for text-to-image, Nano Banana 2.1 with thinking mode enabled scored 1050 points, while the previous generation Nano Banana 2 scored 990 points, and Nano Banana Pro scored 935 points.
In the multi-character consistency benchmark, Nano Banana 2.1 scored 1106, Nano Banana 2 scored 978, and Nano Banana Pro scored 1011, leading the previous generation by nearly 130 points. However, improvement in single-character consistency was smaller: 2.1 scored 1028, the previous generation scored 981, and Pro scored 991.
The data changed the most in the infographic section: Nano Banana 2.1 scored 0.521, Nano Banana 2 only 0.179, and Nano Banana Pro 0.265, about 2.9 times the previous generation. The infographic design score also improved from 961 to 1048.
Below are all the test scores published by Google:
| Review Items | Nano Banana 2.1 (Thinking) | Nano Banana 2.1 (No Thinking) | Nano Banana 2 | Nano Banana Pro |
|---|---|---|---|---|
| Overall Preference for Text-to-Image | 1050 | 1015 | 990 | 935 |
| Infographic Design | 1048 | 1001 | 961 | 912 |
| Infographic accuracy | 0.521 | 0.328 | 0.179 | 0.265 |
| General photo editing | 1026 | 980 | 938 | 939 |
| Single-character consistency | 1028 | 1021 | 981 | 991 |
| Multi-character consistency | 1106 | 1068 | 978 | 1011 |
| Mask/Doodle Retouching | 1049 | 1042 | 965 | 927 |
| Product consistency | 1024 | 981 | 955 | 965 |
| Style transfer | 1062 | 1036 | 991 | 990 |
| Multiple reference image editing | 1066 | 1041 | 988 | 989 |
So how does it compare with competitors? On third-party blind testing platform Arena’s text-to-image leaderboard on October 6, Nano Banana 2.1 ranked 5th with 1328 points, making it the highest-ranked Google model currently. Ahead of it, in order, are two versions of OpenAI’s GPT Image 2.5 (1425 points and 1398 points), GPT Image 2 (1383 points), and Microsoft’s MAI-Image-2.6 (1333 points):

Of course, Nano Banana 2.1 also has its drawbacks.
The official model card also lists the areas where Nano Banana 2.1 currently still falls short, including
- Small text tends to blur, especially at 1K resolution, and long paragraphs and full-page text are still unstable.
- Character consistency between the input and output images isn’t always perfect.
- Mask-based retouching sometimes only follows part of the instructions; during retouching, it sometimes retains the pose of the person in the original image.
- Left and right directions sometimes get confused.
- Advanced capabilities involving world knowledge, 3D reasoning, and factual accuracy also remain limited.
Google says Nano Banana 2.1 will roll out starting October 6 to the Gemini App, AI Mode in Google Search, Google AI Studio, Google Flow, Google Stitch, Google Ads, and the enterprise-focused Gemini Enterprise Platform.
I tested generating images in Gemini, and it has already become Nano Banana 2.1:

On the API side, according to the official Gemini API pricing page, Nano Banana 2.1’s image output price is $30 per million tokens, a direct 50% cut from Nano Banana 2’s $60.
Per image, 1K is US$0.0336, 2K is US$0.0504, and 4K is US$0.0756; using the Batch API cuts that in half again, with 1K costing only US$0.0168. Put simply, generating a 1K image with the standard version costs about NT$1.
| Item (USD) | Nano Banana 2.1 | Nano Banana 2 | Nano Banana 2 Lite | Nano Banana Pro |
|---|---|---|---|---|
| Input (per million tokens) | 1.50 | 0.50 | 0.25 | 2.00 |
| Text and thinking output (per million tokens) | 7.50 | 3.00 | 1.50 | 12.00 |
| 0.5K images (each) | Not provided | 0.045 | Not provided | Not provided |
| 1K image (each) | 0.0336 | 0.067 | 0.0336 | 0.134 |
| 2K images (each) | 0.0504 | 0.101 | Not provided | 0.134 |
| 4K images (each) | 0.0756 | 0.151 | Not provided | 0.24 |
Source: KOCPC Chinese