Not surprisingly, Google officially unveiled the new generation Gemini 3.5 series models at the I/O 2026 event, with the Flash version launching first and the flagship Pro version slated to arrive next month. In the past, the Flash series has been positioned in Google’s product lineup as affordable, fast, but with slightly less capability. However, this time Gemini 3.5 Flash is quite different – it can directly compete with competitors’ Opus 4.7 and GPT-5.5, and has even taken the lead in many benchmarks, which is truly impressive.
More importantly, Gemini 3.5 Flash maintains its consistently high cost-performance ratio, and the default model for Gemini and Google AI Search has been switched to 3.5 Flash. This means even free users can easily experience this next-generation most powerful model, which will certainly put considerable pressure on OpenAI and Anthropic.

Gemini 3.5 Flash Launches: Outperforms Claude Opus 4.7 and GPT-5.5 in Agents and Multimodal, but Still Trails in Long-Form and Reasoning
Google says Gemini 3.5 Flash has made the biggest improvements in “AI agents” and “code development,” calling it their strongest agent and code development model yet.
Compared to its predecessor Gemini 3 Flash, Gemini 3.5 Flash shows significant improvements across nearly all benchmarks released by Google. In the challenging coding and agentic benchmarks, Terminal-bench 2.1 jumped from 58.0% to 76.2%, MCP Atlas rose from 62.0% to 83.6%, and GDPval-AA saw a substantial increase from 1204 to 1656 Elo—an impressive leap overall. What’s more, Gemini 3.5 Flash also outperforms its own previous flagship, Gemini 3.1 Pro, in most benchmarks.
Compared to competitors, in the area of AI agents
- In MCP Atlas tool calling benchmarks, Gemini 3.5 Flash scores 83.6%, Claude Opus 4.7 at 79.1%, and GPT-5.5 at 75.3%, meaning it directly leads the competition by a significant margin.
- In the cross-tool operation Toolathlon benchmark, it also narrowly beat GPT-5.5’s 55.6% with 56.5% (Opus 4.7 was not tested for this item).
Gemini 3.5 Flash also leads in both multimodal benchmarks.
- CharXiv Reasoning (extracting information from complex charts) scores 84.2%, slightly surpassing GPT-5.5 at 84.1% and Opus 4.7 at 82.1%
- MMMU-Pro achieved 83.6%, ahead of GPT-5.5’s 81.2% and Opus 4.7’s 75.2%.
Additionally, in the expert task category, Finance Agent v2 (Financial Analysis and Decision Making) with Gemini 3.5 Flash also achieved 57.9%, surpassing Opus 4.7’s 51.5% and GPT-5.5’s 51.8%.
The following are test items with similar scores
- On Terminal-bench 2.1 (agentic terminal writing), it scored 76.2%, which fell short of GPT-5.5’s 78.2% but outperformed Opus 4.7’s 66.1%.
- OSWorld-Verified (agent computer operation) is 78.4%, nearly identical to GPT-5.5’s 78.7% and Opus 4.7’s 78.0%, with the gap between all three within 1%.
- Blueprint-Bench 2 (agentic spatial reasoning) at 33.6% only trails GPT-5.5’s 36.2% by a little, but has already significantly surpassed Opus 4.7’s 24.5%.
As for coding, while Gemini 3.5 Flash does lead ahead of the previous flagship model Gemini 3.1 Pro, it still falls somewhat behind compared to Opus 4.7 and GPT-5.5:
- SWE-Bench Pro (code generation) scores 55.1%, Claude Opus 4.7 scores 64.3%, lagging behind by about 9%, and GPT-5.5 scores 58.6%.
Complete Test Results Chart:

Even with much stronger performance, Gemini 3.5 Flash doesn’t compromise on speed—it can help complete even lengthy, multi-step tasks in extremely short time, and the cost is often less than half of other comparable models. Looking at this, the arrival of Gemini 3.5 Pro is really something to look forward to.
In addition, those announced at the same event Gemini Spark Their own 24-hour AI agent tool is also now powered by Gemini 3.5 Flash. Google has also upgraded its agent development platform, Antigravity, to version 2.0 with support for the Gemini 3.5 Flash model, now capable of dispatching multiple sub-agents simultaneously to work in parallel and execute multi-step tasks across editors, terminals, and browsers.

Gemini 3.5 Flash is now available and can be used on the following platforms:
- Gemini App
- Google Search’s AI Mode
- Gemini API in Google Antigravity, Google AI Studio, or Android Studio
- Enterprises can access via the Gemini Enterprise Agent Platform and Gemini Enterprise
Gemini 3.5 Pro is currently in use internally at Google and is expected to officially launch next month.
Source: KOCPC Chinese