Earlier, xAI officially released the Grok 4 model. Unsurprisingly, it has once again regained the lead in many benchmark scores, compared to OpenAI、Google …is even more powerful than the latest models, and even calls Grok 4 the smartest in the world. AIAdditionally, to meet the needs of heavy users, xAI also rolled out a new SuperGrok Heavy subscription plan at $300 per month, providing early access to the latest features, such as the upcoming video generation model and AI coding model.

xAI officially unveils the world’s smartest Grok 4 series
At the start of the presentation, xAI and Elon Musk brought up a benchmark called the “Humanities Last Exam” to illustrate how smart Grok-4 is. This exam is extremely challenging; every question is carefully designed by subject-matter experts, with 2,500 in total, covering mathematics, natural sciences, engineering, and all humanities:

Earlier this year, when the “Humanities Ultimate Exam” was first released, most models at the time had accuracy in the single digits. Every question was at the doctoral or even higher research level; even among humans, almost no one could truly answer these questions and score high. Grok-4 achieved postdoctoral level in all fields.
Grok-4 has two model versions, “Grok-4” and “Grok 4 Heavy”. Grok 4 Heavy is the multi-agent version of Grok 4, which uses the “Test and Compute” method to run multiple independent agents simultaneously, then compare their outputs and decide on the best answer.
As can be seen from the figure below, Grok-4 can already solve a quarter of the problems without any tools, achieving 25.4%—higher than Gemini 2.5 Pro’s 21.6% and OpenAI o3’s 21%. With tools, Grok 4 jumps to 38.6%, and Grok 4 Heavy reaches 44.4%. xAI also added that Grok 4 with tools is more powerful and reliable than Grok 3’s “Deep Search” model.

Among other common benchmarks, Grok-4 also performed well, such as on the doctoral-level problem set GPQA, where Grok-4 reached 87.5% and Grok-4 Heavy even scored 88.9%, higher than the strongest models from competitors. In addition, in the AIME American Invitational Mathematics Examination, Grok-4 Heavy achieved a perfect score of 100%. In coding tests, Grok-4 leads other competitors in LiveCodeBench, HMMT, and USAMO:

Grok 4 is smart, but it’s not without its flaws. The rest of Grok 4 still isn’t quite up to par—for instance, it lacks image understanding capabilities. While it does have image generation, that still needs improvement. xAI doesn’t deny this, but has promised that a new version set to be completed within a few weeks will address its visual weaknesses.
As for the API side, Grok 4 also set the best score in the industry on the ARC-AGI benchmark, achieving 15.8% accuracy—double that of the second-place Claude Opus model:
Grok 4 is now live on the Grok platform, but for now it’s limited to paid users; free users can only use Grok 3.
The newly launched SuperGrok Heavy plan at $300 per month offers early access to Grok 4 Heavy and new features. This plan is essentially similar to the ultra-premium subscriptions rolled out by OpenAI, Google, and Anthropic.

Full event video:
Introducing Grok 4, the world's most powerful AI model. Watch the livestream now: https://t.co/59iDX5s2ck
— xAI (@xai) July 10, 2025
Source: KOCPC Chinese