When people are still marveling at Claude Opus 4.6 With OpenAI GPT-5.3-Codex Earlier today, when [their] incredibly powerful abilities Google Suddenly released Gemini 3 Deep Think This major upgrade, specifically designed for complex reasoning tasks in scientific research and engineering domains, achieved astonishing results across multiple authoritative benchmarks, with capabilities surpassing those of Claude and OpenAI’s equivalent models, quickly sparking heated discussion throughout the global tech community.

Gemini 3 Deep Think Sets Multiple Benchmark Records, Significantly Outpacing Competitors
Gemini 3 The most notable aspect of this Deep Think upgrade is its breakthrough performance across multiple challenging benchmarks. In the so-called “Humanity’s Last Exam” (HLE), the model achieved a score of 48.4%, the highest current score without external tool assistance. Even more significant is the ARC-AGI-2 test, a benchmark verified by the ARC Prize Foundation that is considered a key indicator of Artificial General Intelligence (AGI) capabilities. Gemini 3 Deep Think scored an impressive 84.6%, far exceeding the human average of around 60%. In comparison, Anthropic’s Claude Opus 4.6 only reached 68.8%, while OpenAI’s GPT-5.2 came in at 52.9%.

In competitive programming, Gemini 3 Deep Think achieved a 3455 Elo rating on Codeforces, reaching the “Legendary Grandmaster” level, with only 7 human programmers ranking higher globally. Claude Opus 4.6 scored 2352 on Codeforces, with a gap of approximately 47% between the two.
In academic olympiad competitions, the model achieved gold medal level across IMO 2025 (International Mathematical Olympiad), IPhO 2025 (International Physics Olympiad), and IChO 2025 (International Chemistry Olympiad). On the CMT-Benchmark (Advanced Theoretical Physics Test), it scored 50.5%. For MMMU-Pro, it achieved 81.5%, outperforming Claude Opus 4.6 at approximately 75% and GPT-5.2 at approximately 78%.

The only area where it falls slightly short is SWE-Bench Verified (software engineering benchmark), achieving 76.2% on code repair tasks, slightly lower than Claude Opus 4.6’s 80.8% and GPT-5.2’s 80.0%.
A Comprehensive Comparison with Competing Models: How Big Is the Gap Really?
To better understand Deep Think’s market positioning, the following is a complete comparison with its main competitors Claude Opus 4.6 and GPT-5.2:
ARC-AGI-2 (Abstract Reasoning)
Deep Think:84.6% ✅ | Claude 4.6:68.8% | GPT-5.2:52.9%
→ Deep Think outperforms Claude by 23%, outperforms GPT by 60%
Humanity’s Last Exam (Expert-Level Questions)
Deep Think:48.4% ✅ | Claude 4.6:36.7% | GPT-5.2:35.4%
→ Deep Think leads by approximately 32%
Codeforces Elo rating (programming contests)
Deep Think:3455 ✅ | Claude 4.6:2352
Leading by about 47%
MMMU-Pro (Multi-modal Understanding)
Deep Think:81.5% ✅ | Claude 4.6:~75% | GPT-5.2:~78%
GPQA Diamond (Research-Level Q&A)
Deep Think:93.8% ✅ | Claude 4.6:~90% | GPT-5.2:92.4%
SWE-bench Verified (Software Engineering) ⚠️
Deep Think:76.2% | Claude 4.6:80.8% ✅ | GPT-5.2:80.0%
→ Claude leads; Deep Think’s only weakness
AIME 2025 (Pure Mathematics)
Deep Think: 95% (with tools 100%) | Claude 4.6: ~94% | GPT-5.2: 100% ✅
Terminal-Bench 2.0 (Terminal Operations)
Gemini 3 Pro:54.2% | Claude 4.6:65.4% ✅ | GPT-5.2:64.7%
Long context @1M tokens
Gemini 3 Pro:26.3% | Claude 4.6:76% ✅
Based on the above data, Deep Think demonstrates overwhelming advantages in abstract reasoning, expert-level problems, and competitive programming. However, Claude still maintains a clear lead in software engineering, terminal operations, and long-context processing.
Summary of Each Model’s Leading Areas
- Deep Think LeadsARC-AGI-2 (Abstract Reasoning), HLE (Expert Problems), Codeforces (Programming Competition), Science Olympiad, Multimodal Understanding
- Claude LeadsSWE-bench (Software Engineering), Terminal-Bench (Terminal Operations), Long Context Practical Usability, Enterprise Knowledge Work
- GPT-5.2 LeadsAIME (Pure Mathematics), Price Performance
Price Comparison:
- Claude Opus 4.6: $5/$25 per M tokens (most expensive)
- GPT-5.2 Thinking:$1.75/$14 per M tokens
- Gemini 3 Pro: $2/$12 per M tokens (cheapest)
- Google AI Ultra (includes Deep Think): $249.99/month (subscription)
Practical Applications: From Mathematical Review to Crystal Growth Optimization
Google shared multiple real-world application cases in its official blog, demonstrating the value of Gemini 3 Deep Think in real scientific research scenarios. Rutgers University mathematician Lisa Carbone used Deep Think to review mathematical papers, successfully identifying logical errors that were missed during the human peer review process.
Duke University’s Wang Lab used this technology to optimize crystal growth processes, successfully designing thin film formulations exceeding 100 micrometers in thickness. Google’s internal researcher Anupam Pathak also accelerated the design process for physical components using Deep Think. Additionally, the model can directly convert hand-drawn sketches into 3D printing files, demonstrating its multimodal understanding and engineering application capabilities.
Google AI Ultra subscription plan officially launches
Alongside the Deep Think upgrade, Google simultaneously launched a new Google AI Ultra subscription plan, priced at $249.99 per month (approximately NT$7,800), with a 50% discount for the first three months ($125/month). This premium subscription includes full Deep Think functionality, 30TB of cloud storage, Project Mariner (AI agent assistant), Flow (AI video creation tool), Veo 3 (AI video generation model), and other advanced services. Currently, this plan is only available for subscription by users in the United States, while the API interface is open for early applications.
The community reacted enthusiastically, with “Has AGI already arrived?” becoming a hotly debated topic.
This upgrade sparked widespread discussion on social media platforms like X/Twitter. Many users expressed shock at Google’s breakthrough in reasoning, believing Google has temporarily overtaken its competitors. “Does this mean AGI (Artificial General Intelligence) has arrived?” became one of the hottest discussion topics in the tech community.
Google CEO Sundar Pichai, Senior Fellow Jeff Dean, and DeepMind CEO Demis Hassabis all posted on the X platform to celebrate this milestone.
Gemini 3 Deep Think is getting a significant upgrade. We’ve refined Deep Think in close partnership with scientists and researchers to tackle tough, real-world challenges.
And it’s pushing the frontier across the most challenging benchmarks, achieving an unprecedented 84.6% on… pic.twitter.com/5503F4FKcD
— Sundar Pichai (@sundarpichai) February 12, 2026
Perspective
This latest upgrade to Deep Think is truly impressive, especially in abstract reasoning and competitive programming, with a gap from competitors unlike anything we’ve seen before. ARC-AGI-2 leads GPT-5.2 by up to 60%, and Codeforces leads Claude by approximately 47%. These numbers aren’t just statistical differences—they represent a qualitative leap in Deep Think’s complex reasoning chains and multi-step problem-solving capabilities.
However, we must also acknowledge Deep Think’s weaknesses: Claude maintains a clear advantage in software engineering (SWE-bench), terminal operations (Terminal-Bench), and long-context processing. Differences in training objectives and architectural designs across models have pushed AI competition into a new era of specialization—no longer does a single model dominate every domain, but each excels in its own areas.
From a pricing perspective, Gemini 3 Pro’s $2/$12 pricing is highly aggressive, offering Google’s solution strong cost-performance appeal for developers and research institutions with high-volume API call needs. For scientific research, algorithm competitions, and multimodal reasoning needs, Deep Think is undoubtedly the current best choice; however, for software engineering, enterprise knowledge management, and long document processing, Claude remains the more robust option.
Source: KOCPC Chinese