Besides OpenAI,Anthropic Earlier as well, Anthropic released a new version of the Claude Opus 4.6 model, boasting improved coding capabilities. It also marks the first Opus-level model to offer a 1 million token ultra-long context window in its beta version, with more thorough planning and the ability to execute agent tasks for extended periods. Across multiple benchmarks, it achieves state-of-the-art performance in the industry.Claude Subscribers can use it now.

Claude Opus 4.6 Gets Two Major Upgrades: More Stable for Long Tasks, 1M Token Context Window Now Available
The main differences between Claude Opus 4.6 and its predecessor Claude Opus 4.5 lie in improved task quality for long-running tasks and greater stability with large contexts, enabling better planning, sustained agentic task execution, more reliable performance in larger codebases, and stronger code review and debugging capabilities.
Regarding the “1M token context window,” many people may not have a good grasp of what tokens are—you can think of them as the amount of content that can be read and remembered in one go. With a larger capacity, AI can better track multiple documents, long reports, or the broader context of larger coding projects within a single conversation.
Anthropic stated that they found Opus 4.6 automatically focuses on the most difficult parts and quickly processes the simpler parts, makes more mature judgments when encountering ambiguous issues, and maintains high efficiency during extended work sessions.
Opus 4.6 can also be applied to various daily work tasks, such as financial analysis, research, as well as using or creating documents, spreadsheets, and presentations. When paired with Cowork, it can independently complete the tasks you assign.
Now let’s look at the benchmark scores.
In GDPval-AA, a benchmark measuring economically valuable knowledge work such as finance and law, Opus 4.6 outperformed OpenAI’s GPT-5.2 by approximately 144 Elo points, and also surpassed its predecessor Claude Opus 4.5 by 190 points.

Also leading all frontier models on Humanity’s Last Exam (cross-domain complex reasoning benchmark):

In the Vending-Bench 2 test, Opus 4.6 maintained focus for extended periods and earned $3,050.53 more than Opus 4.5.

The following chart shows more benchmark results, outperforming Opus 4.5 in many areas, particularly in Agentic search and Novel problem-solving:

Image source: Claude
Anthropic has also launched Claude in PowerPoint, allowing you to use Claude directly in the PowerPoint sidebar after installation.
In the past, users could already have Claude generate a presentation file, but if they needed to edit it, they still had to manually import it into PowerPoint, which was a bit cumbersome. With the launch of Claude in PowerPoint, users can now generate and edit directly within PowerPoint, receiving ongoing assistance from Claude throughout the creation process.
Source: KOCPC Chinese