Have you ever wondered if the AI summaries that appear at the top of your Google search results could be wrong? The New York Times commissioned AI startup Oumi to conduct large-scale testing, and the results found that Google’s built-in Gemini-powered AI Overview accuracy is only about 90%That means 1 out of every 10 searches gives a wrong answer. With Google’s billions of daily searches, this represents AI Overview Hundreds of thousands of false information items could be generated every hourreaching up to several million [posts] over the course of a single day

From 85% to 91%: Called progress, but error count still alarming
This test uses OpenAI’s SimpleQA BenchmarkA question bank containing over 4,000 verifiable answers used to measure the factual accuracy of generative AI. Oumi began testing during last year’s Gemini 2.5 era, when AI Overview’s SimpleQA score was 85%; when retested this year, the updated AI Overview after Gemini 3 improved to 91% 。
The numbers look like an improvement, but the problem is: even with just a 9% error rate, multiplied by Google’s billions of daily searches, the total number of wrong answers is still astronomical. Oumi’s estimate points out that AI Overview produces a staggering amount of incorrect information every day, reachingmillions of posts。
Error Example: Confidently Citing Incorrect Content
《The New York TimesThe report listed multiple specific error cases, each demonstrating that AI is not “unable to find answers,” but ratherConfidently providing wrong answers even with sources available:
Case 1: Bob Marley Museum Dates
I searched for the date when Bob Marley’s former residence was turned into a museum, and the AI Overview cited three pages: two of them don’t mention any date at all, and the only one that mentions a date, a Wikipedia entry, also listed.Two contradictory yearsDecisively chose the wrong one

Case 2: Yo-Yo Ma Hall of Fame Controversy
I searched for “the date cellist Yo-Yo Ma was selected for the Classical Music Hall of Fame.” The AI Overview cited content from the organization’s official website, yet claimed “this Hall of Fame doesn’t exist at all,” when in fact this Hall of Fame does exist.

The commonality among these errors is that AI is not generating randomly, but ratherSelectively citing incorrect content from erroneous sourcesLies backed by data are more dangerous than complete ignorance.
Google Pushes Back: SimpleQA Has Its Own Issues
Facing this research, Google spokesperson Ned Adriance told The New York Times: “Google believes the SimpleQA benchmark itself contains incorrect information.” He emphasized that Google internally uses another set called SimpleQA Verified The test features questions that have undergone stricter manual review, smaller in scale but higher in quality.
“The study has serious flaws,” Adriance said. “It doesn’t reflect what people actually search for on Google [2].” Google’s position is that SimpleQA tests general knowledge questions rather than the operational or comparative queries users most commonly search for, so SimpleQA results cannot be directly extrapolated to real search behavior.
The limits of benchmark testing itself
However, analysts point out that AI evaluation itself is like an art still in its infancy, with each company adopting its own preferred evaluation methods. The non-randomness of generative AI makes results even harder to reproduce. Even when running the same question again immediately, AI sometimes gets it right and sometimes gets it wrong, making independent verification extremely difficult. Oumi itself even uses AI tools for evaluation, essentially making it “AI judging AI.”
Additionally, SimpleQA’s questions tend to be “simple factual” questions (such as dates, names, and numbers), rather than the action-oriented or comparison-type queries that Google users most commonly search for. The difficulty structure of these two question types is fundamentally different.
The Trust Crisis in the AI Search Era
Regardless of whether this study’s methodology has flaws, a fundamental issue has emerged: when Google places AI Overview at the very top of search results, more and more users are treating AI responses as the definitive answer, rather than—as they used to—treating article links as starting points and interpreting the content themselves. As a result, errors in AI Overview are far more likely to be accepted without question than traditional webpage errors.
Simply put: in the past, Google gave you a page of links, and if you picked wrong, that was your problem. Now Google gives you an answer, and if it’s wrong… whose problem is it?
Source: KOCPC Chinese