New flagship models usually show up with two things: a stack of benchmark charts and a bigger bill. Claude Opus 5.5 arrived on September 22 with the charts, sure, but the bill went the other way.
It costs less per token than the Opus 5 it replaces, and Anthropic says it finishes the same work with fewer tokens, too. (If you're already hunting for the catch, same. There are a couple, and they're hiding in the fine print.)
So I went through everything Anthropic published about it: the launch post, the model docs, the pricing pages, and the Claude Code configuration notes.
The footnotes alone could fill a weekend, and they're where the surprises live, like a new default setting that can change your results before you've touched a line of code.
The short version is right below. After that, I'll get into what changed from Opus 5, how it scores, what it costs, where it runs, and which effort setting to start on before you move real work over.
TL;DR:
Claude Opus 5.5 is Anthropic's newest Opus model, built for long-running agentic coding and knowledge work. Anthropic announced it on September 22, 2026, as the first release in its new Claude 5.5 family.
The headline claim is that it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. That's Anthropic's framing, so treat it as a strong hypothesis to test on your own tasks.
You get a 1M-token context window and up to 128K output tokens, a reliable knowledge cutoff of June 2026, and the API model ID claude-opus-5-5. Adaptive thinking is always on, and you control how hard it thinks with an effort setting.
If the Claude lineup were a restaurant kitchen, Fable 5.1 would be the head chef you call in for the hardest dishes, Sonnet 5.5 the fast line cook, and Opus 5.5 the sous chef who can run most of the service alone.
Anthropic's own model docs now tell developers to start with Opus 5.5 for most workloads. Sonnet 5.5 followed on September 28, and Haiku 5.5 is due in the coming weeks.
If you're still deciding whether Claude fits your work, it's worth comparing a few Claude alternatives first.
Opus 5.5 is cheaper, faster, and more frugal with tokens, and it also changes two defaults that can trip up anyone migrating code. Here's how the two compare on the points that touch your bill and your workflow:
Cost per task dropped more than the price did. Per-token prices fell 20%, but Anthropic says Opus 5.5 also uses fewer tokens to finish the same job, which is how it gets to roughly 40% cheaper on typical workloads.
One early tester used it to audit and fix a 200,000-line codebase in under three hours, a job where Opus 5 took over 20 hours and used 2.5x as many tokens.
The defaults moved. A request that doesn't set an effort level now runs at medium, one notch below Opus 5's high, and thinking can't be disabled. Code that already ran on Opus 5 with thinking on needs no change on that front.
It writes more clearly. Early testers found its writing easier to follow, and the launch post says it puts the most important information up front and follows the writing rules you give it. (If you've ever asked a model for three bullet points and received a short novella, this one's for you.)
Subscribers get more room. Anthropic is raising five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans, and it's giving subscribers a rate limit reset they can save and use whenever they choose.
It behaves better under pressure. Anthropic reports that Opus 5.5 posted the best scores of any model it has tested on its automated behavioral audit, and that it resists prompt injection better than Opus 5.
Anthropic's launch table puts Opus 5.5 ahead of Fable 5.1 and Opus 5 on every test it reports, and ahead of GPT-6 Astra on four of the six where OpenAI has a score. These are Anthropic's numbers, run at max effort (xhigh on Terminal-Bench 4.0):
Astra keeps two wins. GPT-6 Astra leads on AutomationBench, Zapier's test of business workflows across connected apps, and on Terminal-Bench-Science. Zapier counted every safeguard intervention as a failure, which Anthropic says left Opus 5.5 with a lower score than it would get in practice.
Efficiency is the clearer win. At its default medium effort, Opus 5.5 matches Astra on Terminal-Bench 4.0 for about 40% of the cost per task, and on FrontierCode it beats Astra's top score for about a fifth of the cost.
To see how the two do on the same real tasks, and where Astra comes out ahead, read our Opus 5.5 vs Astra comparison.
Opus 5.5 is billed per million tokens (MTok) on the Claude API and cloud platforms. Here are the current list prices next to Opus 5:
The cache-read price is the one to watch. Cache reads make up the majority of agentic and coding work costs, according to the launch post, and that line fell 60%. If your agents reread the same codebase or document set all day, that's where most of your savings will come from.
A few other pricing options are worth knowing about:
If you use Opus 5.5 inside the Claude apps, your plan price covers it, and the plan's usage limits decide how much Opus you get each week.
Opus 5.5 is available to Claude Pro, Max, Team, and Enterprise users. The Free plan doesn't include Opus models, according to Claude's pricing page, so you'll need a paid plan to try it in chat.
Opus 5.5 works in Claude Code from version 2.1.280 or later, so run claude update if you're behind. On the Anthropic API, Amazon Bedrock, Google Cloud, and Claude Platform on AWS, the opus model alias now points to Opus 5.5.
Microsoft Foundry is the one catch, because the opus alias there still resolves to Opus 4.6, so you'll want to select claude-opus-5-5 by its full name. Fast mode works in Claude Code too, if you're the impatient type (no judgment).
If you run Claude Code alongside other AI coding agents, check each tool's model picker, since support rolls out tool by tool. Kiro, for example, now offers Opus 5.5 with experimental support.
Developers can call claude-opus-5-5 on the Claude API today. It's also on Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry, with Bedrock using the ID anthropic.claude-opus-5-5.
Fast mode is the exception. It runs only on Anthropic's first-party Claude API, so you won't find it on Bedrock, Claude Platform on AWS, Google Cloud, or Foundry.
{{templates}}
Effort is the dial that decides how much Opus 5.5 thinks before it answers, and because thinking is always on, it's also your main lever on cost and speed. There are five levels:
The benchmark scores and the default setting don't match. Those headline numbers come from max effort, while the default is medium, so a workflow that never touches effort gets the cheaper, faster setting.
That's less scary than it sounds. Deloitte told Anthropic that Opus 5.5 at its lowest effort caught 72% of known bugs in code reviews, compared with 56% for Opus 5 at high effort, and Rogo reported a similar lowest-effort win on its finance benchmark.
My suggestion is to start on medium, then rerun your most common tasks at low and high and compare the output and the bill. Anthropic's docs recommend the same thing: run a fresh effort sweep on your own evals whenever you move up from an older model.
Anthropic pitches Opus 5.5 as a daily driver for serious coding and knowledge work, and the early-tester examples in the launch post back that up with specifics. Keep in mind that these are vendor-published results.
This is the use case Anthropic leans on hardest. It says Opus 5.5 is particularly good at long, sprawling jobs like codebase-wide migrations and audits, and one early tester finished a 680,000-line code migration in less than a day.
In an internal test, Anthropic had Opus 5.5 and Fable 5.1 translate HAProxy from C into Rust. Both rewrites passed nearly all of HAProxy's regression tests, but Opus 5.5 finished in 9.5 hours compared with 12 and cost 51% less.
Opus 5.5 is built to keep working for hours with little supervision. A developer at Clio handed it a task spanning six repositories, let it run overnight, and said it stayed on task for over 18 hours.
Column's engineering team says it delegates to subagents far more effectively, and Anthropic reports it's much less likely than recent models to take hard-to-reverse actions. That combination matters if you run AI agents unattended.
Anthropic asked Opus 5.5, Fable 5.1, and Opus 5 to write reports on a company's quarterly results from a copy of the web, and a grader failed any invented figure or quote. Sixteen of Opus 5.5's 18 reports passed, while neither of the other models passed once.
It also leads GDPval-AA v2.1, an Artificial Analysis test of real-world work across 44 occupations, by more than 100 Elo points over Fable 5.1.
According to Anthropic's docs, Opus 5.5 reads dense charts, diagrams, and screenshots much more precisely than earlier models, and Anthropic calls it its best Opus model for vision and computer use. That helps with document extraction and multi-step tasks that hop between apps.
Few model launches are all upside, and Opus 5.5 has a few catches worth knowing before you swap model IDs:
{{cta}}
Fable 5.1 still sits above Opus 5.5 in Anthropic's lineup, at $10 input and $50 output per MTok, which is 2.5x the price. Anthropic's advice is to start with Opus 5.5 and move up to Fable 5.1 for demanding reasoning or when Opus 5.5 at higher effort still falls short.
Opus 5.5 outscores Fable 5.1 on every benchmark in Anthropic's launch table, though Anthropic says the difference in its own use is smaller than the scores suggest.
For most teams, the practical order is to try Opus 5.5 first. Our Opus 5.5 vs Fable 5.1 comparison covers the cases where Fable 5.1 is still worth paying for.
Sonnet 5.5 arrived on September 28, 2026, six days after Opus 5.5, at half the price: $2 input and $10 output per MTok, with the same $0.20 cache reads. Anthropic pitches it for well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets.
It even beats Opus 5.5 on Terminal-Bench 4.0 (70.6% vs 66.4%). Anthropic still says Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment," so a sensible split is Sonnet 5.5 for quick, scoped work and Opus 5.5 for long agent runs.
Claude Opus 5.5 is a rare upgrade where the newer model is also the cheaper one, so for most teams on Opus 5 the real work is picking the right effort setting, which you can do in a single afternoon.
Pick a handful of tasks your team runs every week, like a code review, a bug fix, a research brief, and a spreadsheet model, then run each one on Opus 5.5 at low, medium, and high. Compare the quality and the token bill side by side, and pin the lowest setting that still passes.
If a lot of your team's AI work runs through an AI teammate like Lindy, which lets you pick the model for each task, check which models your workspace lists before you plan the switch around it.
Once the test is done, you'll know which setting to use, and if Anthropic's 40% figure holds on your workload, the savings show up on every invoice after that.
Claude Opus 5.5 handles long-running agentic coding, research, and knowledge work, including codebase migrations, code review, financial models, reports, and multi-step computer use. It has a 1M-token context window and can stay on a single task for hours.
Claude Opus 5.5 beats GPT-6 Astra on most benchmarks in Anthropic's launch table, including Terminal-Bench 4.0 (66.4% vs 57.9%), while Astra leads on AutomationBench and Terminal-Bench-Science. These are Anthropic's numbers, so test both if you're choosing between Claude and ChatGPT.
Opus 5.5 is available in Claude Code from version 2.1.280. On the Anthropic API, Amazon Bedrock, Google Cloud, and Claude Platform on AWS, the opus alias selects it automatically, and Microsoft Foundry users select claude-opus-5-5 by name.
Most Opus 5 users should upgrade to Opus 5.5, because it's cheaper per token, generates output over 30% faster, and uses fewer tokens per task.
Before you switch API code, remove any setting that disables thinking or forces tool use, move computer use agents to the new toolset, and set your effort level explicitly.
Opus 5.5 is the stronger model for complex, open-ended work, while Sonnet 5.5 costs half as much at $2 input and $10 output per MTok and responds faster. Both have a 1M-token context window, so Sonnet 5.5 suits well-scoped everyday tasks and Opus 5.5 suits long, judgment-heavy jobs.
