OpenAI The long-rumored next-generation large language model (LLM) “GPT-5“It might really be coming.” Recently, on WebDev Arena (a well-known platform that compares LLM model performance using double-blind methods), an LLM suspected to be codenamed “Openclaw” quietly appeared. And this mysterious contestant with the codename “Openclaw” unexpectedly surpassed the current strongest model, Grok-4, in all tests, leaving AI practitioners and relevant figures around the world exclaiming: “GPT-5, here it comes!”

“Openclaw” Bursts onto the Scene: Anonymous Review Reveals GPT-5’s Debut
WebDev Arena is a website specifically designed to evaluate the capabilities of large language models, employing unified prompts and a double-blind scoring mechanism to avoid any brand bias. Recently, a user named “Lisan al Gaib” discovered an exceptionally outstanding model on this platform, whose performance far surpasses Grok-4. According to the comparison charts shared by this user, the model codenamed “Openclaw” demonstrates remarkable results in front-end page generation, creating neural network animations with artistic visuals, smooth animations, and deep immersion.
GPT-5 vs Grok-4
exact same prompts, wildly different output in web lmarena https://t.co/2ER6cNg5HD pic.twitter.com/N4hJ4L5eAV
— Lisan al Gaib (@scaling01) July 25, 2025
The prompt used is as follows:
「Create a stunning, interactive animation of a neural network or brain-like graph structure—use artistic colors, smooth transitions, and beautiful visuals. The page should feel alive, immersive, and impressive, with no buttons—just scrolling or continuous animation. Make it breathtaking.」(創建一個令人驚嘆的神經網絡或類腦圖結構的互動式動畫——使用藝術性的色彩、平滑的過渡和精美的視覺效果。整個頁面應充滿活力、沉浸感和震撼力,不使用任何按鈕,僅透過滾動或持續的動畫進行互動。讓其令人叹為觀止。)
This task places extremely high demands on frontend code generation—perfect execution requires not only understanding natural language but also integrating animation, graphics, and interactive logic. Openclaw’s output quality has stunned the entire community: this could well be GPT-5’s first real-world debut.
Suspected multiple versions revealed simultaneously: “Openclaw”, “Nectarine”, “Starfish”
More notably, in addition to “Openclaw,” other codename information for the GPT-5 series has emerged in the community, seemingly representing its different tier versions:
-
GPT-5 Standard Edition: Openclaw
-
GPT-5-mini Lite Version: Nectarine
-
GPT-5-nano Micro Edition: Starfish
– they are all reasoning models
– Openclaw > Nectarine > Starfish on my testing with LisanBench
– Nectarine and Openclaw beat previous high-score of 62 with the starting word “camping”:
o1 – 62
Sonnet 4 Thinking – 49
o3 – 47
GPT-4.5 – 44Openclaw – 81
— Lisan al Gaib (@scaling01) July 25, 2025
This naming convention also seems to continue OpenAI’s past internal codename strategy for model versions, and the community generally speculates that these are all disguised codenames used to conceal official names during the development and testing phase.
Reddit community and developers reveal more clues.
According to observations from the Reddit developer community, OpenAI has recently been quietly redirecting some “o3” requests to a new model for processing. Even on anonymous evaluation platforms like LMArena, OpenAI has registered another model under the identity “zenith.” Some developers have found that this model can handle extremely difficult math problems that GPT-4 or o3 previously could barely solve, and its logical style is clearly different from before.


These signs all point to one possibility: OpenAI is extensively testing the boundaries of GPT-5’s capabilities through multiple anonymous test models.
Internal staff and invited users get early access: the feedback is stunning.
According to partOnline community members revealedSome employees of companies in non-technical fields have already been granted access to the GPT-5 preview version. Although the specific identities of these invited companies cannot be confirmed due to signed non-disclosure agreements, this indirectly confirms that GPT-5 has now entered a broad internal testing phase.

OpenAI CEO Sam Altman also hinted in a recent interview that his initial experience with GPT-5 has been “very powerful,” and he has begun subtly revealing the model’s performance in various settings.
Sam Altman on GPT 5:
“ GPT-5 is the smartest thing. GPT-5 is smarter than us in almost every way. You know, and yet here we are. ”
This might be the last podcast before the big release! pic.twitter.com/MgfSHMRjGk
— Chris (@chatgpt21) July 23, 2025
Programming capabilities fully evolved, possibly Claude Sonnet 4
From first-hand feedback provided by trial users, GPT-5’s most highly praised capability is “coding.” According to reports, GPT-5 not only performs exceptionally well on competitive programming problems, but more importantly, its ability to handle real-world engineering scenarios has seen a qualitative leap. Specifically, even when faced with massive codebases containing large amounts of legacy “spaghetti code,” GPT-5 can effectively modify and optimize them, demonstrating unprecedented comprehension and refactoring abilities. This has even overshadowed Claude Sonnet 4, which was once hailed as the “King of Coding” in developer circles.
This change could have a major impact on OpenAI’s market strategy, especially in the high-revenue AI coding assistant market. For example, Cursor relies on Claude models to provide powerful coding assistance and has already surpassed $100 million in annual revenue. OpenAI clearly hopes to win back this slice of the pie (especially since OpenAI’s recent attempt to acquire Windsurf reportedly fell through).

It is understood that one of the core development goals of GPT-5 is to integrate the traditional GPT models with the o series (models known for their reasoning capabilities) into a unified interface. According to user feedback, GPT-5 can automatically adjust its reasoning ability based on question difficulty: simple spelling problems automatically trigger a low-resource mode, while complex questions such as “optimizing a database architecture that hasn’t been maintained for 10 years” activate a deep reasoning process.
Reassuring the market and investors: OpenAI hasn’t hit a wall—it’s overtaking on the curve.
In 2024, there were claims that large model development had hit a “wall,” with the argument that the marginal benefits of pretraining were declining. But the arrival of GPT-5 seems to refute this view (assuming those tests really were all GPT-5). Through smarter reasoning strategies and post-training techniques such as reinforcement learning and instruction fine-tuning, OpenAI’s latest GPT-5 shows that this path hasn’t reached a bottleneck—it’s simply shifting toward a more efficient way to reach the next stage. For AI hardware companies like Nvidia and for investors, GPT-5’s future performance could help steady the ship, and it also provides greater confidence for future data center expansion and AI applications.
Is GPT-5 really coming?
In summary, GPT-5 has already demonstrated epoch-making progress in terms of technical integration, reasoning strategies, and application capabilities alike. It not only fills the gaps OpenAI had in programming and engineering applications, but may also redefine the path to achieving AGI. Although the specific release date and scope of availability have not yet been announced, judging from community reactions, trial feedback, and statements from leadership, GPT-5 is undoubtedly on the brink of release, waiting only for the final push—the official launch.
Source: KOCPC Chinese