Skip to main content
Models & Research

Introducing Claude Opus 5.5 - Anthropic

We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

By Precis Daily Newsroom8 min read1,689 words
Illustration for: Introducing Claude Opus 5.5 - Anthropic
Illustration
Key points
  • We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
  • Claude Opus 5.5 is our first release since we called for pacing the frontier .
  • On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date.

We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. Claude Opus 5.5 is our first release since we called for pacing the frontier . It was tested before release by external evaluators, including Frontier Design and METR . On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models. Here are some of the improvements you can expect from Opus 5.5: Performance. Opus 5.5 is a major step up from Opus 5. It’s the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish. Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it’s been given, and it’s more resistant than Opus 5 to prompt injection. We’ve also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card . Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program , and verified cybersecurity practitioners will be able to use Opus 5.5 for their work. Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5. In addition to the price drop, we’re increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. We’re also providing subscription users a rate limit reset, which you can now save and use whenever you choose. Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one. Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety. On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest. Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol Agentic coding Terminal-Bench 4.0¹ Agentic coding Terminal-Bench 4.0¹ 66.4% 55.8% 52.3% 57.9% 37.3% Agentic coding FrontierCode v1.1 (Main) Agentic coding FrontierCode v1.1 (Main) 54.4% 50.3% 48.0% 53.3% 47.5% Agentic coding CursorBench 4.0 Agentic coding CursorBench 4.0 57.8% 51.8% 46.6% — 41.7% Knowledge work GDPval-AA v2.1 Knowledge work GDPval-AA v2.1 1846 1735 1708 1542 1588 Business workflows AutomationBench² Business workflows AutomationBench² 40.0% 31.4% 26.9% 41.4% 28.8% Multidisciplinary reasoning Humanity's Last Exam Multidisciplinary reasoning Humanity's Last Exam 67.7% with tools 65.6% with tools 63.6% with tools 57.2% with tools — Agentic scientific research Terminal-Bench-Science 0.1³ Agentic scientific research Terminal-Bench-Science 0.1³ 58.7% 52.6% 29.0% 64.6% 22.4% Computer use OSWorld 2.1 Computer use OSWorld 2.1 81.8% partial 80.7% partial 74.0% partial — — Visual chart recognition Chartography Visual chart recognition Chartography 89.0% with tools 88.4% with tools 83.4% with tools — — Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI; these represent each model’s highest score. Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5’s performance on these benchmarks. 1 Terminal-Bench 4.0: The standard error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the other Claude models. The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI. 2 AutomationBench: AutomationBench results were run and reported by Zapier. These runs were performed without fallback models, so safeguard interventions were considered failures—this resulted in a lower score than Claude Opus 5.5 would achieve in practice. Claude Opus 5.5 results come from Zapier’s own evaluation during early access. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier’s public leaderboard. 3 Terminal-Bench-Science 0.1: The standard error is ±3.5–5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, within noise. The GPT-6 Astra figure is as reported by OpenAI. Where Opus 5.5’s advantage is very clear is efficiency. It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs. Prices per 1M tokens Claude Opus 5.5 Claude Opus 5 Cache reads $0.20 $0.50 Input tokens $4 $5 Output tokens $20 $25 Cache writes $5 $6.25 Fast mode for Opus 5.5 is also available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens. Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less. Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, while on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost. Agentic terminal coding Agentic coding: FrontierCode Agentic coding: CursorBench Agentic terminal coding Agentic coding: FrontierCode Agentic coding: CursorBench Terminal-Bench 4.0 Accuracy vs Cost Opus 5.5 0 10 20 30 40 50 60 70 Score (%) 2 5 10 20 Cost per attempt (USD, log scale) Low Med High Xhigh Max Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command line interface. Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. It matches GPT-6 Astra at about 40% of the cost. FrontierCode v1.1, main set Accuracy vs Cost Opus 5.5 35 40 45 50 55 0 Score (%) 0.50 1 2 5 10 Cost per task (USD, log scale) Low Med High Xhigh Max FrontierCode measures whether an agent’s code changes would be merged. At default effort (medium), Opus 5.5 scores 54.6%, higher than all other models, beating GPT-6 Astra’s top score (53.3%) for about a fifth of the cost per task. CursorBench 4.0 Accuracy vs Cost Opus 5.5 25 30 35 40 45 50 55 60 0 Score (%) 1 2 5 10 20 Cost per task (USD, log scale) Low Med High Xhigh Max CursorBench evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions. At default effort (medium), Opus 5.5 scores 52.5%, compared to 51.8% for Fable 5.1 (max) and 46.6% for Opus 5 (max). It beats GPT-5.6 Sol’s top score (41.7%) by 11 points for about a third of the cost per task. Our early testers reported similar efficiency and intelligence gains: GitHub Clio Lovable Quantium Spotify Optiver Column Kiro GitHub Clio Lovable Quantium Spotify Optiver Column Kiro Quote “Developers want agents that can take on real software work and finish it. In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable.” Author Mario Rodriguez, Chief Product Officer

Sources

Summarized from the linked originals.

Related stories

Illustration for: Anthropic's mid-tier Claude climbs the rankings
Models & Research

AI Anthropic's mid-tier Claude climbs the rankings PLUS: Pick the right Claude model with one quick test Good morning, AI enthusiasts, and welcome to our 5,342 new readers. OpenAI takes the stage today for one of its most hyped days of the year.

The Rundown AI7 min