Introducing Claude Sonnet 5.5 - Anthropic
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.

- Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
- It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.
- Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5.
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work. Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design. Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks. Performance. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5’s 10.3%. It scores two points below Opus 5.5 on GDPval-AA, a test of real-world work across a variety of occupations. And it’s strong on long-horizon work and image understanding—it’s the first Sonnet model to beat Pokémon Red working only from screenshots. Collaboration. Like Opus 5.5, Sonnet 5.5 writes more clearly than our previous generation of models; early testers described it as a better partner for collaboration than Sonnet 5. Its speed also makes it well suited to fast iteration on less complex tasks. Cost. Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task than its predecessor. Speed. Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date. Alignment and safety. On our automated behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment. Because its cybersecurity capabilities are comparable to Opus 5’s, it’s the first Sonnet model to launch with cyber safeguards and fallbacks like those we’ve developed for our most capable models. Its biology safeguards are the same as Sonnet 5’s. Both safeguards target a narrow set of high-risk requests; routine software development and most life sciences work are unaffected. Sonnet 5.5 improves on Sonnet 5 across domains—in some cases dramatically. On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol Agentic coding Terminal-Bench 4.0 Agentic coding Terminal-Bench 4.0 70.6% 10.3% 66.4%¹ — Agentic coding FrontierCode 1.1 (Main) Agentic coding FrontierCode 1.1 (Main) 46.2% Max² 42.4% 54.4% 49.3% 52.1% Xhigh Agentic coding CursorBench 4.0 Agentic coding CursorBench 4.0 55.5% 34.1% 57.8% — Knowledge work GDPval-AA v2.1³ Knowledge work GDPval-AA v2.1³ 1844 1449 1846 1487⁴ Knowledge work AA-Briefcase v1.1³ Knowledge work AA-Briefcase v1.1³ 1811 1359 1822 1483⁴ Multidisciplinary reasoning Humanity’s Last Exam Multidisciplinary reasoning Humanity’s Last Exam 64.5% with tools 54.9% with tools 67.7% with tools — Computer use OSWorld 2.1 Computer use OSWorld 2.1 80.1% partial 57.0% partial 81.8% partial — Visual chart recognition Chartography Visual chart recognition Chartography 61.6% no tools 15.6% no tools 64.4% no tools 53.6%⁴ no tools For details on how we run our evaluations, see the Sonnet 5.5 System Card . The charts below plot each model’s score against its cost per task at every effort level. As effort goes up, models typically work for longer, leading to a higher cost per task but generally also a higher score. The closer a point is to the top left of the chart, the more capability it delivers per dollar. On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score for about a tenth of the cost per task. It complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost. Agentic terminal coding Agentic coding: FrontierCode Agentic coding: CursorBench Knowledge work: AA-Briefcase Agentic terminal coding Agentic coding: FrontierCode Agentic coding: CursorBench Knowledge work: AA-Briefcase Terminal-Bench 4.0 Accuracy vs. cost Sonnet 5.5 0 10 20 30 40 50 60 70 Score (%) 1 2 5 10 Cost per attempt (USD, log scale) Low Med High Xhigh Max Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command-line interface. At Medium effort, the default in the Claude apps, Sonnet 5.5 far exceeds Sonnet 5’s best score for less than a tenth of the cost per task. Terminal-Bench and OpenAI did not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here. FrontierCode v1.1, main set Accuracy vs. cost Sonnet 5.5 30 35 40 45 50 55 0 Score (%) 0.25 0.50 1 2 5 10 20 Cost per task (USD, log scale) Low Med High Xhigh Max FrontierCode measures whether an agent’s code changes would be merged. At High effort, the default on the Claude Platform, Sonnet 5.5 matches GPT-6 Sol’s best score for about a fifth of the cost per task.² CursorBench 4.0 Accuracy vs. cost Sonnet 5.5 20 30 40 50 60 0 Score (%) 0.50 1 2 5 10 Cost per task (USD, log scale) Low Med High Xhigh Max CursorBench evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions. Sonnet 5.5 at Low effort exceeds Sonnet 5’s best score for less than a tenth of the cost per task. CursorBench 4.0 does not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here. AA-Briefcase v1.1 Accuracy vs. cost Sonnet 5.5 900 1100 1300 1500 1700 1900 0 Elo 0.10 0.20 0.50 1 2 5 10 20 Cost per task (USD, log scale) Low Med High Xhigh Max On AA-Briefcase, a new benchmark of long-horizon knowledge work, Sonnet 5.5 at Medium effort bests Sonnet 5’s best score for about one ninth of the cost per task.³ Sonnet 5.5’s jump in performance is particularly noticeable in coding. At High effort on FrontierCode, it scores 10 points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task. On CursorBench, which tests models on tasks from real Cursor coding sessions, its best score is within about two points of Opus 5.5. Early testers appreciated how quickly Sonnet 5.5 can understand a codebase. They were also struck by its efficiency: in head-to-head runs, it batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs. Epic Games Every CodeRabbit SpaceXAI Base44 Unity Creator Epic Games Every CodeRabbit SpaceXAI Base44 Unity Creator Quote “In Epic’s early testing, Claude Sonnet 5.5 cleared the same quality bar you’d expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting.” Author Daniel Vogel, Chief Operating Officer
Sources
Related stories

Add secure Web Search to Claude Desktop with Amazon Bedrock AgentCore
Claude Desktop on Amazon Bedrock provides powerful AI assistance, but without integrated web search, responses are limited to the model’s training knowledge…

Last Week in AI #345 - 5 new models, 9 misalignment incidents, some Dots
Top NewsAnthropic and OpenAI race to release smarter and cheaper modelsSources:Anthropic launches Claude Opus 5.5 with stricter safeguards for…

GLM-5.3 and the spread of advanced cyber capabilities - Anthropic
GLM-5.3 and the spread of advanced cyber capabilities Cole McFaul, Robert Xiao, Tripp Gallagher Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits.

Anthropic's mid-tier Claude climbs the rankings
AI Anthropic's mid-tier Claude climbs the rankings PLUS: Pick the right Claude model with one quick test Good morning, AI enthusiasts, and welcome to our 5,342 new readers. OpenAI takes the stage today for one of its most hyped days of the year.