• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - Gemini Returns to Glory? Google Launches Gemini 4 Argon Flagship Model, Single-Output Limit Soars to 1 Million Tokens, Tops 13 of 19 Benchmarks

Gemini Returns to Glory? Google Launches Gemini 4 Argon Flagship Model, Single-Output Limit Soars to 1 Million Tokens, Tops 13 of 19 Benchmarks

Rocky by Rocky
October 1, 2026
in AI Trends and Related News

It’s been quite a while since Gemini released a flagship model that made people sit up and take notice. The last time Google brought out its top-tier model was Gemini 3.1 Pro Preview in February this year. At Google I/O in May, although the company had teased that Gemini 3.5 Pro would arrive the following month, it ultimately never launched; instead, the lightweight Flash series kept updating from 3.5, 3.6, and 3.7 all the way to 3.8. This has also left many people wondering: is Gemini almost done for?

Just a short while ago, Google finally unveiled its all-new generation Gemini 4, starting with the flagship model Gemini 4 Argon. According to the officially released test results, Argon took first place in 13 out of 19 benchmarks, while its maximum single output limit jumped from 64,000 tokens to 1 million tokens. Although the initial rollout is not yet directly available to general users, at least the debut of Gemini 4 Argon shows that the Gemini flagship model, after a period of silence, is finally starting to make moves.

Google launches Gemini 4 Argon: up to 1 million tokens per output, aimed at coding, legal and financial work, and cybersecurity defense.

Earlier, Google announced on its official blog the launch of Gemini 4 Argon A new generation of frontier models—the most advanced tier among all models. Gemini 4 Argon focuses on three areas: real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.

The features can be roughly divided into 4:

  1. Single output limit raised to 1 million tokens
  2. The ability to program and perform enterprise knowledge work.
  3. visual understanding tasks such as reading charts and watching long videos
  4. The ability to identify and patch cybersecurity vulnerabilities.

First, the extremely long output limit. Gemini 4 Argon can output up to 1 million tokens in a single response, up from 64,000 tokens in the previous generation, which is arguably the highest in the industry. By comparison, competitors’ output limits are mostly 128,000 tokens.

Google says that when a model has enough room to think deeply and generate hundreds of thousands of tokens in a single run, it is more likely to solve a difficult problem in one go, rather than being cut off halfway through its reasoning.

On the programming and enterprise knowledge work front, Google says its internal engineers have already put Argon to use in their daily work, from routine debugging and large-scale code migrations to algorithm design. In addition, it specifically highlights Argon’s performance in white-collar work such as finance, law, and tax, including conducting multi-step financial research, drafting legal documents, or running an entire business process from start to finish on automation platforms like Zapier.

Argon is also particularly strong at knowledge work that requires “visual understanding,” such as analyzing professional charts, finding specific details in very long videos, or taking action directly after reading through a series of documents.

Cybersecurity defense is the area Google focused on the most this time.

Google mentioned that they deliberately trained Gemini 4 Argon to be a model skilled in cybersecurity defense, able to find software vulnerabilities on its own, verify whether those vulnerabilities actually exist, and then automatically write patches. In early testing, Argon found a serious vulnerability in medical software used by hospitals around the world that could expose sensitive personal data, which other frontier models had not previously discovered.

This time, the official also released a comparison table of 19 tests, comparing against OpenAI’s GPT-6 Astra, as well as Anthropic’s Claude Fable 5.1 and Claude Opus 5.5. Argon took first place in 13 of them, tied for first in 1, and trailed in 5. The full results table is as follows:

Category Test item Gemini 4 Argon GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5
Knowledge work Vals Index 68.9% 63.1% 65.8% 67.0%
AutomationBench 51.3% 41.4% 31.4% 42.5%
Vals Finance Agent v2 65.4% 53.5% 58.9% 58.6%
Harvey Legal Agent Test 19.6% 5.4% 6.7% 3.8%
Agentic Coding DeepSWE v1.1 77.9% 74.1% 67.4% 74.2%
FrontierSWE v2 55.0% 65.5% 56.3% 62.3%
Vibe Code Bench 91.9% 89.6% 90.3% 90.3%
Terminal-bench 4.0 57.4% 58.2% 57.9% 66.4%
Machine learning engineering PostTrainBench 45.3% 44.3% 40.2% 49.3%
Science and Mathematics Terminal-Bench Science 0.1 57.6% 68.1% 52.6% 63.3%
LABBench 2 88.8% 85.4% 68.6% 73.1%
RiemannBench 76.0% 72.0% 65.6% 69.6%
Long context GraphWalks (within 128K tokens) 99.7% 98.7% 91.4% 90.6%
GraphWalks (256,000 to 1,000,000 tokens) 84.2% 71.8% 65.0% 66.8%
Computer operation Agent’s Last Exam 39.5% 34.2% — 38.2%
OSWorld-2.0 69.2% 72.6% — —
Multimodal understanding Chartography 71.6% 71.0% 46.2% 66.3%
LVBench 91.7% 87.5% 79.7% 83.7%
Information security CWE-bench v1 68.0% 68.0% 58.0% 67.0%

First, look at where Argon leads: the largest gaps are concentrated in enterprise knowledge work.

On the Vals Index, which measures AI’s economic value in finance, programming, legal, and tax work, Argon scores 68.9%, Claude Opus 5.5 67.0%, Claude Fable 5.1 65.8%, and GPT-6 Astra 63.1%—not a very large lead:

The larger gap was with AutomationBench, introduced by Zapier. This test examines whether a model can automate a business process from start to finish. Argon scored 51.3%, while second-place Claude Opus 5.5 scored 42.5% and GPT-6 Astra scored 41.4%, a lead of nearly 9%:


The most absurd part is the legal agent test by legal AI company Harvey. This test is extremely difficult, and all four models scored very low, but Argon scored 19.6%, while GPT-6 Astra scored only 5.4% and Claude Opus 5.5 only 3.8%. Argon’s score is more than three times that of its competitors:

In terms of coding, which many people care about, in DeepSWE v1.1 testing, Argon scored 77.9%, Claude Opus 5.5 scored 74.2%, and GPT-6 Astra scored 74.1%, which Google says is the current record. Vibe Code Bench also narrowly beat the two Claude models’ 90.3% with 91.9%:

It also performs well in science and visual understanding. For example, on LVBench, which tests long-video understanding, Argon scored 91.7%, GPT-6 Astra 87.5%, and Claude Opus 5.5 83.7%. However, on Agent’s Last Exam for computer operation, Argon’s 39.5% is only slightly higher than Claude Opus 5.5’s 38.2%.

Of course, Argon doesn’t win in every category either. It trails in 5 tests, and in 2 of them Argon is even last among the four models.

On the more difficult software engineering benchmark FrontierSWE v2, Argon scored only 55.0%, while GPT-6 Astra scored 65.5% and Claude Opus 5.5 scored 62.3%. On Terminal-bench 4.0, Argon scored 57.4%, Claude Opus 5.5 led by a wide margin at 66.4%, and GPT-6 Astra also scored 58.2%.

Unfortunately, Google did not include the results for Gemini 3.1 Pro or Gemini 3.8 Flash this time, so it is not possible to directly see how much Gemini 4 Argon has improved over the previous generation in various tests.

At this stage, Gemini 4 Argon is only available through Google DeepMind’s Fairwind program and is not yet open to general users. However, Google has said it will collect feedback from early testers while adjusting safeguards, then make it available to developers, enterprises, and general consumers as soon as possible; the first group will be paid API customers and Google AI Ultra subscribers.

API pricing has now been announced, split into two phases: introductory pricing and standard pricing, but it doesn’t say how long the introductory period will last:

Model (per million tokens) Input Cache input Output
Gemini 4 Argon (introductory price) 2 US dollars USD 0.10 10 US dollars
Gemini 4 Argon (official price) 4 US dollars 0.20 US dollars 20 US dollars
Claude Opus 5.5 4 dollars 0.20 USD 20 US dollars
GPT-6 Astra 10 US dollars 1 US dollar 50 US dollars

Source: KOCPC Chinese

Tags: aiArtificial IntelligenceGeminiGemini 4Gemini 4 ArgonGoogle

Recent Posts

  • Google Play adds usage-based billing for AI apps.
  • Google Gemini will discontinue the Gems system, turning your most-used prompts into skills.
  • Gemini Returns to Glory? Google Launches Gemini 4 Argon Flagship Model, Single-Output Limit Soars to 1 Million Tokens, Tops 13 of 19 Benchmarks
  • Volvo EX90 launches in Taiwan starting at NT$2.95 million: 800V architecture, 254 TOPS compute; for the first month, the Plus trim adds NT$30,000 and includes ventilated seats.
  • AI Misidentifies Car! Flock Safety License Plate Recognition Error Leads to 23-Year-Old Woman’s 13-Day Wrongful Imprisonment; Senate Holds Hearing to Investigate AI Surveillance Flaws

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology