Just a few days ago, Google proudly unveiled Gemini 3.5 Flash at I/O 2026, claiming that this model with “Pro-level reasoning, Flash-level latency” Vals AI Finance Agent v2 In the financial agent benchmark test, using 57.9% Its performance beat GPT-5.5 (51.8%) and Claude Sonnet 4.6 to take first place. Google’s official data is written in black and white, showing that Gemini 3.5 Flash leads all competitors in financial agent tasks.

However, just a few hours after the good news spread throughout the community, X users… Chetaslua(@chetaslua) posted three screenshots: he brought in three top-tier models and asked them the same addition problem, 300 + 140. As it turned out, Gemini 3.5 Flash actually answered it wrong with complete confidence.
Gemini 3.5 Flash vs GPT-5.5 instant vs Sonnet 4.6
Remember guys. #1 in Finance Agent v2. SOTA performance right here. lol 🤣
Prompt : ” 300+140=460 Is this correct? Breakdown? ” https://t.co/zvNXMjlyKE pic.twitter.com/La8siKwjWP
— Chetaslua (@chetaslua) May 22, 2026
The AI finance champion even got 300+140 wrong, so we also tried.ReproduceConfirmed that Gemini 3.5 Flash does give incorrect answers, just like the earlier question about which is larger, 0.9 or 0.11.

The King of Finance
First, let’s see how the term “champion” came about.Vals AI Finance Agent v2 It is currently the industry’s most authoritative financial AI agent benchmark, simulating complex tasks in real-world financial workflows: financial report analysis, market data interpretation, multi-step financial decision-making, and more.
According to the comparison data published by Vals AI:Gemini 3.5 Flash:57.9% 🥇/GPT-5.5: 51.8%/Claude Opus 4.7: 51.5%/Claude Sonnet 4.6: 51.0%/Gemini 3.1 Pro: 43.0%, beating second-place GPT 5.5 by 6%.

A 6-percentage-point lead is no small margin. Vals AI’s official account also confirmed on X: “Google’s Gemini 3.5 Flash took first place on the Finance Agent benchmark v2, beating GPT-5.5 by 6 points.” Gemini 3.5 Flash is indeed the strongest financial AI agent model today, at least on this benchmark.
Then someone asked it, “300+140=?”
Chetaslua called Gemini 3.5 Flash, GPT-5.5 instant, and Claude Sonnet 4.6 to the front and asked them the same question:
300 + 140 = 460, right? Please explain step by step.
Gemini 3.5 Flash:
「Yes, that is completely correct!」(對,完全正確!)
百位數:300+100=400
十位數:0+40=40
個位數:0+0=0
加總:400+40+0=460

That’s right — not only did it get the answer wrong, but it also provided a seemingly well-reasoned derivation to prove that 300+140 equals 460. If you only look at the derivation steps and not the result, each line is correct on its own, but put together they are wrong. 400+40=440, not 460.

GPT-5.5 instant:
「Yes. 300+140=460… 推導:300+100=400,400+40=440。所以正確答案是 300+140=440。不是 460。」

It starts with “yes,” then contradicts itself. It blindly agrees at first, only realizes something is off during the reasoning process, and finally makes a last-minute correction. Slightly better than Gemini—at least it gets the right answer in the end—but if you only read the first sentence and leave, you’ve already been misled.
Claude Sonnet 4.6:
「No, that’s incorrect. 300+140=440, not 460。」
300+100=400
400+40=440

Correct, concise, and confident from start to finish. No fluff.
Where’s the problem?
This is not a case-specific failure of Gemini 3.5 Flash, but a structural problem with the current AI evaluation system. In Google’s own positioning, Gemini 3.5 Flash’s core selling point is agentic tasks: it scored 83.6% on MCP Atlas (the highest in the industry), 56.5% on Toolathlon, and 78.4% on OSWorld-Verified. These scores are real and represent genuine significant progress in complex tool calling and multi-step task execution. But there is a crucial logical leap between being an “agentic task expert” and being a “finance expert”:Strong performance on agent tasks =/= strong financial ability; strong financial benchmarks =/= strong arithmetic ability.

What does Finance Agent v2 test? Reading financial reports, interpreting market data, and executing multi-step financial workflows—these tasks rely more on semantic understanding, long-text reasoning, and instruction following. In other words, Gemini 3.5 Flash scoring high on financial benchmarks doesn’t mean it truly “understands numbers.”
The “math” of large language models is essentially pattern matching, not logical computation. When the model sees “300+140=460, right?”, there may be numerous fragments like “XXX+YYY=460” in its training data, plus the sycophancy bias that “users asking ‘right’ usually expect a positive answer,” leading it to choose the wrong path. What’s more embarrassing is that Gemini 3.5 Flash not only answered incorrectly but also fabricated a derivation process to make the error appear correct—this is a typical symptom of large model hallucination.。
Operationalization Challenges in AI Benchmarking
This incident has once again sparked debate within the AI community over the “credibility” of benchmark testing. On Reddit r/singularity, the most popular discussion…MessageWhat is it?
「Google 精心挑選的基準測試,剛好避開了所有 Flash 會出糗的題目。」
This statement reveals the unwritten rule of today’s AI industry:Whoever releases the benchmark can choose questions that favor themselves. Every AI company will advertise its model as “SOTA” on the dimensions where it performs best, but it won’t proactively tell you that the same model can even get simple addition wrong.
In fact, if you look closely at the official datasheet released by Google, Gemini 3.5 Flash does not stand out on benchmarks that require genuine logical reasoning, such as Humanity’s Last Exam (40.2%, lower than 3.1 Pro’s 44.4%) and ARC-AGI-2 (72.1%, lower than 3.1 Pro’s 77.1%). Its strength lies in agentic tasks, not pure reasoning or math.
This went viral online not just because Gemini got a math problem wrong, but because The label “financial AI agent” itself implies that this is a model good at handling numbers.Financial statement analysis requires numbers, market forecasting requires numbers, and asset pricing requires numbers. If a model claims to be the best in finance yet cannot even handle the most basic elementary-school addition, how can people trust it to process complex numbers in real-world financial scenarios?
Actually, opinions on Gemini 3.5 Flash have been quite polarized since its launch — its benchmark scores are quite high, but many netizens’ real-world tests have left them less than impressed with its performance. It looks like Google may still need some time to fine-tune it and figure out where the problem lies.
打开了 AntiGravity 测试一下 Gemini 3.5 Flash,一个明显的感觉:
速度快到令人发指…
但就是狂输出代码但不解决问题…
按照完成任务所用时间来算,还不如 Opus 4.6 的 20%,甚至还不如 Gemini 3.1 Pro!
服了…大家可以继续嘲讽美国豆包了…
妈的垃圾! https://t.co/YSriPbiHPP pic.twitter.com/3RRzIgoKrg
— Crypto_Painter (@CryptoPainter) May 19, 2026
Source: KOCPC Chinese