As Anthropic launched last week Claude Opus 5.5 Later, unsurprisingly, the second model in the Claude 5.5 family, Sonnet 5.5, also officially went live earlier.
Sonnet 5.5 continues the capability upgrades of the Claude 5.5 family. It not only responds more than 30% faster than the previous-generation Sonnet 5, but also uses fewer tokens for the same work, saving up to 30% per task. In addition, its coding and computer-use abilities have improved significantly, approaching Opus 5.5 on multiple benchmarks. Here’s a summary for everyone.

Anthropic launches Claude Sonnet 5.5: 30% faster output, up to 30% lower cost per task, and surpasses Opus 5.5 on Terminal-Bench 4.0.
According to Anthropic Introduction: Claude Sonnet 5.5 has 5 main improvements this time:
- Faster and more economical.: Output speed is over 30% faster than Sonnet 5, making it the fastest Sonnet currently. API pricing hasn’t changed—still $2 per million input tokens and $10 per million output tokens—but completing the same task uses fewer tokens and tool calls, and in official tests the cost per task can be up to 30% lower.
- Code and AI Agent CapabilitiesThe ability to write code, understand codebases, and execute multi-step tasks in the command line has improved significantly, making it especially suitable for day-to-day development and bug fixing.
- Computer Operation and Chart ComprehensionThis is one of the areas where it has improved the most over the previous generation: it can watch the screen, use the mouse and keyboard on its own to complete tasks, and read charts. It is also the first Sonnet model to beat Pokémon Red using only game screenshots, and it has stronger long-duration task capabilities, having handled work lasting several hours in early tests.
- Writing and DesignIt inherits some of Opus 5.5’s upgrades in natural writing, and when building interfaces it also fixes more details, making documents, presentations, and spreadsheets more polished.
- Safety ProtectionSonnet 5.5’s cybersecurity capabilities are significantly improved over Sonnet 5, having reached a level close to Opus 5, making it the first Sonnet model to introduce advanced cybersecurity protections and a fallback mechanism.
Since it costs only half as much as Opus, many people must be wondering: can Sonnet 5.5 directly replace Opus 5.5?
Most everyday tasks are fine, but for harder programming tasks, Opus 5.5 is still the preferred and more reliable choice.
In other words, for well-scoped development, bug fixes, documentation, and presentations, switching to Sonnet 5.5 delivers similar results at half the cost. But for planning large-scale architecture and handling the most complex programming problems, it’s still best to leave those to Opus 5.5. That’s exactly how Creator’s creative programmer Kevin Ngo uses them: let Opus 5.5 define the architecture first, then hand implementation over to Sonnet 5.5.
Looking at the benchmark scores makes it clearer.
In OSWorld 2.1, which tests computer operation, Sonnet 5.5 got 80.1%, while Opus 5.5 got 81.8%; on Cursor’s CursorBench 4.0, it was 55.5% versus 57.8%; on GDPval-AA v2.1, which simulates workplace knowledge work, it was 1,844 points versus 1,846 points, with the two models’ scores extremely close. Even more notable, in the Terminal-Bench 4.0 test for completing multi-step tasks in the command line, Sonnet 5.5 also got 70.6%, surpassing Opus 5.5’s 66.4%:
| Test items | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (command-line agent tasks) | 70.6% | 10.3% | 66.4% | — |
| FrontierCode 1.1 (Code) | 46.2% | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 (Code) | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 (Workplace Knowledge Work) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (Workplace Knowledge Work) | 1811 | 1359 | 1822 | 1483 |
| Humanity’s Last Exam (Available Tools) | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 (computer use) | 80.1% | 57.0% | 81.8% | — |
| Chartography (chart understanding, no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
For the hardest coding task, FrontierCode 1.1, Sonnet 5.5 scored 46.2%, trailing Opus 5.5’s 54.4% by 8% and also losing to GPT-6 Sol’s 49.3%. Officials also added that Sonnet 5.5’s FrontierCode score with Max enabled is actually lower than with Xhigh, because it runs extra code reviews on its own, resulting in timeouts or changes outside the task scope:

On the security front, Sonnet 5.5’s cybersecurity capabilities are significantly improved over Sonnet 5, so Anthropic has for the first time added cybersecurity protections to a Sonnet model similar to those in Opus 5.5.
This won’t affect everyday development such as routine bug fixing, but higher-risk cybersecurity requests will automatically be answered by Sonnet 5 instead. In addition, Sonnet 5.5 also has an anti-distillation classifier design, used to prevent others from extracting its reasoning process to train their own models.
Claude Sonnet 5.5 is now fully live across all Claude App, Claude Code, Claude Platform, AWS, Google Cloud, and Microsoft Azure. The API model name is claude-sonnet-5-5.
API pricing is $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5 and only half that of Opus 5.5:
| Per million tokens (USD) | Sonnet 5.5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|
| Input | 2 | 4 | 2 |
| Output | 10 | 20 | 10 |
| Cache read | 0.2 | 0.2 | 0.2 |
| cache write | 2.5 | 5 | 2.5 |
As for the Haiku 5.5 model in the Claude 5.5 family, officials said it will be released in the coming weeks.
Source: KOCPC Chinese