As expected, just like beforeRumorsUnlike previous models, GPT-5.6 wasn’t immediately rolled out to the masses upon release. OpenAI just officially previewed the new GPT-5.6 series, with GPT-5.6 Sol being the first to debut. Compared to GPT-5.5, its agent task capabilities have been elevated to the next level, and its security is on par with Mythos Preview. Initially, it will only be available to a select group of trusted partners before gradually expanding access.

GPT-5.6 Sol introduces max/ultra reasoning modes, with upgrades across both security and biological capabilities
The naming of this GPT-5.6 series has some changes compared to the past.
OpenAI saidStarting from GPT-5.6, the numbers represent model generations, while Sol, Terra, and Luna represent fixed capability tiers. Sol is the most powerful flagship model; Terra is the version that balances everyday tasks with cost; Luna is the fastest and most affordable version.
This naming scheme is more intuitive than the old ones like mini, nano, pro, or Thinking – at least regular users can grasp it faster. Sol is the most powerful, Terra is the middle tier, and Luna is the affordable fast option.
First to debut is the GPT-5.6 Sol model, offering two reasoning modes:
- Max reasoning effort mode: Allows Sol to spend more time processing complex problems.
- Ultra mode: More like the AI agent workflows people are familiar with today, it doesn’t rely on a single agent to complete tasks, but instead uses subagents to accelerate complex work. Simply put, Ultra breaks down tasks and assigns them to multiple AIs for parallel processing, then consolidates the results.
In the capability testing section, GPT-5.6 Sol showed the most significant improvements in “using commands to operate computers to complete complex tasks” and “handling exploit-related work.”
First is TerminalBench 2.1, which tests whether models can autonomously plan, input commands, read error messages, correct their direction, and ultimately complete the task in a command-line environment.
GPT-5.6 Sol Ultra scored 91.9%, GPT-5.6 Sol came in at 88.8%, slightly beating Claude Mythos 5’s 88.0%. GPT-5.6 Terra hit 84.3%, on par with Claude Fable 5, and even the cheapest GPT-5.6 Luna reached 82.5%. This means it’s not just Sol that got stronger this time—Terra and Luna have already caught up to or surpassed the GPT-5.5 level.

Next up is GeneBench v1, a test that can quickly show how many output tokens are spent to produce how many results.
The official charts show that GPT-5.6 Sol’s curve reaches over 30%, significantly higher than GPT-5.5’s peak of around 23%. Moreover, Sol achieves these higher scores without requiring nearly as many output tokens as Terra. While Terra can also approach around 30%, it needs nearly 50K output tokens to do so. Luna, on the other hand, is the low-cost option, with scores landing in the 10% to 15% range.

ExploitBench is for testing cybersecurity capabilities. As can be clearly seen from the chart, Sol’s success rate curve can reach over 70%, approaching Mythos Preview’s position, while using significantly fewer output tokens than Mythos Preview.

In comparison, GPT-5.5 stalls at just over 40%, while Terra pushes past 50%, and Luna comes in at around 30%. This is also why OpenAI says Sol approaches Mythos Preview but produces only about one-third the output tokens. It’s not just about higher scores—efficiency has clearly improved as well.
Lastly, ExploitGym also tests whether AI can autonomously discover vulnerabilities and complete exploitation, showing a similar trend.
GPT-5.6 Sol’s 6-hour curve reaches above 30%, and the 2-hour curve is also significantly higher than GPT-5.5. Terra can also reach above 20% with extended time, but Luna requires a very large number of output tokens to slowly climb up:

In other words, GPT-5.6 Sol is not only better at finding vulnerabilities, but also more capable of sustaining progress in long-duration, multi-step vulnerability research workflows.
GPT-5.6 Sol 更擅長協助使用者發現並修復漏洞,而不是可靠地執行端到端攻擊。隨著這些能力持續進步,我們的優先事項是確保這些能力能到達並幫助防禦者,讓他們用來找出弱點、開發修補程式,並更廣泛地強化系統。 — OpenAI
To prevent GPT-5.6 from being used for malicious cyber attacks, GPT-5.6 also employs a multi-layered safety mechanism. The model itself refuses to assist with illicit cyber attack requests, and the system performs real-time checks on high-risk responses during content generation, pausing generation when necessary and handing it off to a larger reasoning model for further review.
If a policy violation is detected, the content will be blocked before it reaches the user. Additionally, OpenAI also combines account-level risk signals with continuous monitoring to reduce the likelihood of model misuse.
Finally, on the availability front, GPT-5.6 is not yet fully available. OpenAI says that during the preview period, the GPT-5.6 model will initially be accessible via the API and Codex for a select group of trusted partners and organizations, before gradually rolling out to more ChatGPT, Codex, and API users.
OpenAI also stated in its announcement that although it conducted a limited preview with a small group of trusted partners this time at the request of the US government, it does not believe this government priority access process should become the default practice long-term, as it would prevent users, developers, enterprises, cybersecurity defenders, and global partners who need advanced AI tools from gaining timely access to the latest capabilities.
Source: KOCPC Chinese