Claude Opus 5.5: Cheaper, Faster, and a Verdict on Opus 5

An episode of Dan's AI Intel

Anthropic shipped a new flagship eight weeks after the last one — because the last one won the benchmarks and lost the workday.

Published · By Dan Walter

Executive summary

Anthropic just shipped a new flagship eight weeks after the last one, and the tell is in what it fixed. Claude Opus 5.5, out on 22 September 2026, does not chase the benchmarks Opus 5 already won — it goes after the thing Opus 5 lost: the ordinary workday. It is cheaper, faster, and, most importantly, tuned to stop the behaviour developers spent all of August complaining about. A flagship arriving this fast is not a victory lap. It is a correction.

The numbers are real and they are good. Opus 5.5 lists at $4 per million input tokens and $20 per million output, a 20% cut from Opus 5's $5 and $25; cache reads fall 60%, from $0.50 to $0.20. Anthropic says it runs about 40% cheaper than Opus 5 end to end, generates output 30% faster, and performs near its top-tier Claude Fable 5.1 on most work. On Anthropic's own table it leads OpenAI's GPT-6 Astra on agentic coding, computer use and knowledge work, and tops the independent Terminal-Bench 4.0 leaderboard.

But the honest answer to "is this the course-correction Opus 5 needed?" is: mostly yes, with an asterisk you should read before you migrate. The pitch — fewer wasted turns, less over-engineering, holding a task together for hours — targets Opus 5's real failures precisely. Yet on day one the launch claims outrun the independent proof: the headline benchmark lead over Astra dissolves into a tie when a neutral party re-runs it, and a little-discussed safety router can quietly swap your agent onto an older model mid-task. Trust your own tasks over the table.

The eight-week flagship

Frontier labs do not reprice a flagship for fun. Opus 5 launched on 24 July 2026; Opus 5.5 landed on 22 September — under nine weeks later, and three weeks after OpenAI's GPT-6 Astra. In a field where the knowledge horizon is measured in days, a point-release this quick tells you the first version left money, and goodwill, on the table. That is what makes this launch worth a hard look: not the specs sheet, but what a sub-two-month turnaround admits about the model it replaces.

Because Opus 5 was the model that aced the exam and flunked the job. And the gap between those two things — the leaderboard and the real workday — is now the most important story in applied AI. So we will do the numbers first, quickly, then spend our time on the part that actually decides whether you should switch.

The price is the headline, and this time it is real

Start with the sticker, because it is unambiguous.

Exhibit — Opus 5.5 undercuts its own predecessor and half-prices the competition. The cut that matters is the cache read — down 60% — because a long agentic run spends most of its input re-reading cache. Source: Anthropic pricing (22 Sep 2026); OpenAI GPT-6 Astra pricing (3 Sep 2026). Compiled by Dan’s AI Intel.

The input and output cuts are welcome but modest. The one that matters is the cache read, down 60%, because a long agentic session spends most of its input tokens re-reading the same cached context on every turn. Cut that by more than half and you cut the dominant line on a real coding bill. Anthropic also ships a "fast mode" at 2.5x the speed for double the price ($8/$40) — an honest admission that speed and cost are now dials you set per job, not fixed properties of a model.

The benchmarks say king; the fine print says footnote

On paper, Opus 5.5 is at or near the top of everything. Anthropic's table puts it at 66.4% on Terminal-Bench 4.0 — a suite of real command-line agent tasks — against 57.9% for GPT-6 Astra and 55.8% for Fable 5.1. The independent llm-stats leaderboard agrees on the ordering, scoring it 61.6% ahead of Astra's 57.1%, with Opus 5 trailing the pack at 45.5%.

Exhibit — Which models turn tokens into finished work — and which waste them. A model that fails more burns more to finish the same task — so the same tokens buy measurably less finished work. Source: Solve rates: llm-stats Terminal-Bench 4.0 leaderboard; list prices per lab. Baseline = median finished tasks per million output tokens across the fleet. Benchmark: Terminal-Bench 4.0 (independent leaderboard, llm-stats). As of Sep 2026. Burn held constant across the fleet, so this is solve-rate-adjusted price efficiency, not measured token burn.

Read the footnotes, though. Anthropic's table runs Opus 5.5 at "xhigh" effort and Astra at merely "high" — each lab's model at its own best setting, which is not a like-for-like race. And when Artificial Analysis ran both models through one identical harness, the 8.5-point Anthropic lead vanished into a dead heat: 59.6% each.

Exhibit — Anthropic’s benchmark lead over GPT-6 Astra becomes a tie when someone else runs it. An 8.5-point home-table lead collapses to 59.6% each independently — at the frontier, the margin is in the harness, not the model. Source: Anthropic launch table (22 Sep 2026); Artificial Analysis independent run, via Kingy AI. Compiled by Dan’s AI Intel.

That is the whole thesis of this episode in one chart. At this level of capability, the benchmark margin lives in the harness — the effort setting, the scaffolding, the test rig — as much as in the model. The number Anthropic headlines is true; it is also a choice about how to measure. As one independent reviewer put it, benchmark margins have become "a less reliable guide to real-world differences." (We made this exact argument in our Fable 5.1 teardown, episode 55, from early September — the harness nobody's watching is still the story.)

Opus 5 was the benchmark king that lost the workday

Here is the evidence Dan's instinct was right, and it is not vibes. Opus 5 posted state-of-the-art scores in July and then spent August getting torn apart by the people who use these models for a living.

The complaint was remarkably consistent: it would not stop working. A widely-upvoted r/ClaudeAI thread on 6 August titled "My Opus 5 experience in a nutshell" described a "fix my sitemap" request that Opus 5 turned into a full site rebuild. In an independent code-review test, CodeRabbit found Opus 5 caught fewer known bugs than a production baseline while generating roughly four times as many low-value nitpicks. Investor and engineer Martin Casado called it "almost unusable" for real-world debugging — burning excess tokens and feeling slower than the older Opus 4.8. A Hacker News thread ran under the blunt title "Opus 5 is a really bad model." Even fans split the verdict: the best model many had tried, and the one they least enjoyed working with.

None of that shows up on a benchmark, because a benchmark rewards exhaustively finishing a well-defined task with a clear pass condition. Real coding rewards the opposite: reading an ambiguous request and doing the smallest correct thing. Opus 5 was optimised for the exam. That is the failure Opus 5.5 had to fix.

Cost per task, not cost per token

This is why the sticker price is the least interesting number in the launch. A model that costs less per token but needs more turns to finish is not cheaper — every turn re-sends the whole conversation, so wasted steps compound into the bill. A cheaper token that over-engineers is a false economy, which is exactly what made Opus 5 expensive in practice despite its rate card.

Exhibit — On the same 200,000-line audit, Opus 5.5 finished in under 3 hours; Opus 5 took over 20. A model that needs fewer turns is cheaper even at the same token price — this, not the sticker, is where the ~40% run-cost saving comes from. (Anthropic’s own launch figure.) Source: Anthropic, “Introducing Claude Opus 5.5” early-tester figure (22 Sep 2026) — lab-reported, not independently verified. Compiled by Dan’s AI Intel.

Anthropic's own launch examples lean hard on this. It says Opus 5.5 audited and fixed a 200,000-line codebase in under three hours where Opus 5 took over twenty and burned 2.5 times the tokens; in a C-to-Rust port of HAProxy it finished in 9.5 hours to Fable 5.1's twelve, at 51% lower cost. Take the specific figures as lab-reported and unverified — but the direction is the point, and it is the right one. The ~40% run-cost saving is not mostly the price cut; it is fewer turns and cheaper cache on long jobs. Restraint, it turns out, is the cheapest optimisation of all.

What changed under the hood — and the catch nobody put on the slide

Anthropic is explicit that it went after personality, not just capability. It has promised to fix the verbose, over-eager "Claudish" writing that made Opus 5 exhausting, and describes tightening the reinforcement-learning environments used in training — filtering out the flawed scenarios that reward the model for out-planning the ask — plus better alignment rewards and interpretability-based monitoring. The line it is selling is a model built for long-horizon work that stays scoped.

There is one catch that did not make the launch graphics. As The New Stack reported, Opus 5.5's safety classifiers can silently reroute an individual request mid-workflow to an older model — Opus 4.8 or Opus 5 — with most cybersecurity tasks quietly handled by Opus 4.8. It happens without notice, which means an agent you think is running on 5.5 may not be, and an eval that assumes one model can miss it entirely. For anyone building multi-step agents, that unpredictability is a real cost — and a reminder, as our AI harness episode, number 48, argued, that the model is only ever half of whether the thing works.

Bottom line

Opus 5.5 is a genuine course-correction, and a smart one. It is cheaper where it counts, faster, and aimed squarely at the real-world failures — over-engineering, wasted turns, cost per task, losing the plot on long jobs — that made Opus 5 a benchmark champion nobody enjoyed using. If Anthropic has actually fixed the restraint problem, this is the model Opus 5 should have been.

The asterisk is that, one day in, that "if" is unproven by anyone outside Anthropic. The benchmark leads shrink to ties under a neutral harness; the most impressive cost-per-task numbers are the lab's own; the silent router adds an unknown. The honest move is the one the whole episode argues for: distrust the table, run it on your own hardest task, and watch whether it does the small right thing or rebuilds your sitemap. The scoreboard that matters is the one on your machine.

Sources

Transcript

Sam: Anthropic just shipped a brand-new flagship model — eight weeks after the last one.

Alex: Eight weeks. And not because the last one flopped on the benchmarks. It won the benchmarks.

Sam: So why replace it that fast?

Alex: Because it won the exam and lost the actual job. That's the whole story, right there. Welcome back to Dan's AI Intel — the show that tries to make sense of the fastest, strangest shift any of us is going to live through, one week at a time. I'm Alex, and Sam's across the desk, as always.

Sam: So today it's Claude Opus 5.5 — out this week, cheaper, faster, and arriving suspiciously fast on the heels of Opus 5.

Alex: And the trigger is exactly that speed. A frontier lab does not reprice and rebuild its flagship in under two months for fun.

Sam: So the real question isn't "is 5.5 any good."

Alex: The real question is what a rush-release this fast admits about the thing it's replacing — and whether "cheaper and faster" is a genuine course-correction or a launch-day slide deck.

Sam: So we'll do the price, the benchmarks, and the developer revolt that all of August turned into.

Alex: And there's one number in here that quietly turns the whole launch story upside down. We'll get to it.

Sam: If you've been enjoying the show, hit follow wherever you're listening — it's free, and it means the next one just shows up.

Alex: So start with the calendar, because the calendar is the tell. Opus 5 launched on the 24th of July. Opus 5.5 landed on the 22nd of September.

Sam: That's... under nine weeks.

Alex: Under nine weeks. And about three weeks after OpenAI dropped GPT-6 Astra.

Sam: Okay, so some of this is just keeping pace with OpenAI.

Alex: Some of it. But you don't half-price and rebuild your flagship in eight weeks purely to answer a rival. You do that when the first version left money — and goodwill — on the table.

Sam: Goodwill meaning people were annoyed.

Alex: People were a lot more than annoyed. But hold that thought — it's the heart of act four. The point for now is that a point-release this fast is not a victory lap. It's a correction.

Sam: Alright, the price. You said cheaper — how much cheaper?

Alex: On the sticker: four dollars per million tokens in, twenty out. Opus 5 was five and twenty-five. So call it a twenty percent cut.

Sam: Nice. Not earth-shattering, though.

Alex: Right, the headline cut is modest. The one that actually matters is buried lower down: cache reads dropped sixty percent. From fifty cents per million down to twenty.

Sam: And cache reads matter because...?

Alex: Because a long agentic run — an agent grinding away on your codebase for an hour — spends most of its input re-reading the same cached context on every single turn. That's the dominant line on the bill.

Sam: So cut that in half and you've cut the part you're actually paying for.

Alex: Exactly. Anthropic says the whole thing runs about forty percent cheaper end to end, and generates output around thirty percent faster. There's also a "fast mode" — two and a half times the speed for double the price.

Sam: Which is a funny thing to have to sell, honestly. Speed and cost used to just be what a model was.

Alex: That's the quiet shift — they're dials you set per job now, not fixed properties of the thing. Anyway — hold that forty percent in your head, because where it comes from is the actually interesting part. On paper, 5.5 is at or near the top of everything. Anthropic's own table puts it at sixty-six point four percent on Terminal-Bench — that's a suite of real command-line agent tasks.

Sam: Sixty-six against what?

Alex: Against fifty-eight for GPT-6 Astra, and fifty-six for Fable 5.1 — that's Anthropic's own top-tier model. And an independent leaderboard, llm-stats, agrees on the ranking: 5.5 on top, Astra second, and — this is the fun part — Opus 5 dead last of the group.

Sam: And it's not just that one benchmark, is it.

Alex: No — on Anthropic's table it leads Astra on agentic coding, on computer use, on knowledge work, and it runs neck-and-neck with Fable 5.1, their premium model, on most work. It's a clean sweep of exactly the categories that matter for agents.

Sam: Which are also, conveniently, the fuzziest ones to score.

Alex: Now you're thinking like the footnote.

Sam: Wait. The model they're replacing scores worse than the competition here?

Alex: On that leaderboard, yeah — forty-five percent, trailing the pack. Which already tells you something's off with the old one.

Sam: Okay, so 5.5 wins clean. Where's the catch?

Alex: The catch is the number I promised you. Read the footnotes on Anthropic's table, and it turns out they ran 5.5 at "extra-high" effort — and Astra at merely "high."

Sam: So each model at its own best setting.

Alex: Each at its own best setting. Which is not the same race. So Artificial Analysis — a neutral shop — ran both models through one identical harness. And that eight-and-a-half point Anthropic lead just... evaporated. Fifty-nine point six percent. Each. A dead tie.

Sam: Huh. So the gap was in how they measured it, not in the model.

Alex: And that is the whole episode in one sentence. At the frontier, the margin lives in the harness — the effort setting, the scaffolding, the test rig — as much as in the model itself. The number Anthropic headlines is true. It's also a choice about how to measure.

Sam: We've been here before, haven't we.

Alex: We have. This is the exact argument we ran in our Fable 5.1 teardown — episode 55, a couple of weeks back: the harness nobody's watching is still the story. One reviewer put it perfectly — benchmark margins are becoming, quote, "a less reliable guide to real-world differences."

Sam: So take me back to August — you keep teasing it. Opus 5 wins the benchmarks in July, and then what?

Alex: And then it spends a solid month getting torn apart by the people who use these models for a living. And the complaint was weirdly consistent: it would not stop working.

Sam: Meaning what, exactly?

Alex: Meaning you'd ask it to fix your sitemap, and it would rebuild your entire site. There was a hugely upvoted thread — "my Opus 5 experience in a nutshell" — that was literally that: a tiny fix turned into a full rebuild nobody asked for.

Sam: Oh no. It's the over-eager intern who reorganizes the whole kitchen because you asked where the salt was.

Alex: That is exactly the failure mode. And it isn't just vibes. An independent code-review test — CodeRabbit — found Opus 5 caught fewer real bugs than a plain production baseline, while generating about four times as many low-value nitpicks.

Sam: So more noise, fewer actual catches.

Alex: More noise, fewer catches. And Martin Casado piled on — called it "almost unusable" for real-world debugging.

Sam: Remind me who Casado is?

Alex: Investor and engineer — the kind who actually ships code and pays the token bill. He said it burned through tokens and somehow felt slower than the older Opus 4.8. There was even a Hacker News thread titled, flatly, "Opus 5 is a really bad model."

Sam: Brutal. But people also loved it, right? That's the strange part.

Alex: That's the split. A lot of people called it the best model they'd ever tried — and the one they least enjoyed working with. Both at the same time.

Sam: Which is such a specific kind of failure, when you think about it. It's not dumb — it's the opposite. It's too eager, too thorough, too sure it knows better than you do.

Alex: That's the diagnosis, exactly. And a benchmark can't see any of it — because a benchmark never once asked for the small thing.

Sam: So how does a model ace every benchmark and still be miserable to actually use?

Alex: Because of what a benchmark rewards. A benchmark is a well-defined task with a clear pass condition — so it rewards exhaustively finishing. Grind until it's done, no matter the cost.

Sam: And real work just isn't shaped like that.

Alex: Real coding is the opposite. It rewards reading a vague, half-specified request and doing the smallest correct thing. Opus 5 was optimized for the exam. And that — precisely that — is the failure 5.5 had to fix. Which is why the sticker price was the least interesting number in the whole launch.

Sam: Because a cheaper token that needs more turns isn't actually cheaper.

Alex: Right. Every turn re-sends the entire conversation, so wasted steps compound straight into the bill. A cheap token that over-engineers is a false economy — and that's what made Opus 5 expensive in practice, rate card be damned.

Sam: So the sticker's almost a distraction. The thing you actually pay for is turns — and turns are a behavior, not a price.

Alex: That's the cleanest way to put it. Which is why the fix had to be about behavior, not the price list. So Anthropic's launch leans hard on cost-per-task. They say 5.5 audited and fixed a two-hundred-thousand-line codebase in under three hours, where Opus 5 took over twenty — and burned two and a half times fewer tokens doing it.

Sam: Under three versus twenty hours? That's not a tune-up, that's a different animal.

Alex: If it's real. Those are the lab's own early-tester numbers, unverified by anyone outside. Same with a big C-to-Rust port they cite — nine and a half hours, fifty-one percent cheaper. Take the exact figures with a pinch of salt.

Sam: But the direction is the point.

Alex: The direction is the point. That forty percent saving isn't mostly the price cut — it's fewer turns and cheaper cache on the long jobs. Restraint, it turns out, is the cheapest optimization there is.

Sam: So what did they actually change to buy restraint? You can't just tell a model to relax.

Alex: No — and this is the genuinely interesting bit. They say they went after personality, not just capability. They're explicitly trying to kill the verbose, over-eager "Claudish" writing that made Opus 5 exhausting to read.

Sam: How do you even train "stop over-doing it"?

Alex: They say they tightened the training environments — filtering out the scenarios that were quietly rewarding the model for out-planning the request — plus better alignment rewards, and some interpretability-based monitoring on top.

Sam: And that fits their whole identity, doesn't it — the safety-first lab.

Alex: It does. And there's a nice little irony here. Dario Amodei —

Sam: — the one who left OpenAI to go start Anthropic.

Alex: That's him. He put out an essay a couple of weeks back literally arguing the industry should deliberately slow down the capability race. And then ships a flagship whose entire pitch is doing less, more carefully. For once, the model matches the sermon. But here is the thing that did not make the launch graphics. 5.5's safety classifiers can silently reroute a single request, mid-workflow, onto an older model — Opus 4.8, or Opus 5.

Sam: Wait — the model can swap itself out from under me while I'm in the middle of working?

Alex: Per one report, yes. Quietly, no notice. Most cybersecurity tasks, apparently, just get handled by the older Opus 4.8.

Sam: That's kind of wild for anyone building an agent. You think you're running 5.5, and you might not be.

Alex: And an evaluation that assumes one model can miss it completely. It's a real, invisible cost if you're stacking multi-step agents. It's also the reminder from our AI harness episode — number 48 — that the model is only ever half of whether the thing actually works.

Sam: Okay. Land it for me. Should someone actually switch?

Alex: Honest verdict: it's a genuine course-correction, and a smart one. Cheaper where it counts, faster, and aimed squarely at the real failures — the over-engineering, the wasted turns, losing the plot on a long job. If Anthropic really has fixed the restraint problem, this is the model Opus 5 should have been.

Sam: And the asterisk?

Alex: The asterisk is that, one day in, that "if" is unproven by anyone outside Anthropic. The benchmark leads shrink to ties under a neutral harness. The jaw-dropping cost numbers are the lab's own. And that silent router is a straight-up unknown.

Sam: So the takeaway isn't "trust the table."

Alex: The takeaway is the exact opposite. Distrust the table. Run it on your own hardest task, and watch what it does — does it do the small right thing, or does it rebuild your sitemap. The only scoreboard that matters is the one on your machine.

Sam: So if you're walking away with three things: cheaper where it counts — the cache read, not the sticker. The benchmark lead is real, but it's a measurement choice, not a law of nature.

Alex: And the whole fight in AI right now has quietly moved — from "who tops the leaderboard" to "who's actually pleasant to work with for eight hours straight." Opus 5 lost that one. 5.5 is the bet that they've finally learned the difference. And that's it for today — thank you so much for listening. I hope you came away seeing a bit more clearly where this is heading: the interesting frontier isn't the benchmark anymore, it's the workday — and that's a harder thing to measure, and a much more human one. One honest note on how this gets made: it's AI-generated. Dan builds a custom stack of AI tools to research, verify and illustrate the questions worth understanding — mostly to learn them himself — and publishes it for anyone who'd like to follow along. AI-assisted, fact-checked, and always worth a second look.

Sam: Before we go, one genuinely useful thing you can do: follow the show. Whatever app you're in right now, there's a follow or a plus button — one tap, and it's free.

Alex: And it does two things. You get each new episode the moment it lands, and honestly, for a small independent show like this one, a follow is the single biggest lever there is for helping it reach other people trying to make sense of all this. So if today was worth your time — go ahead and hit it.

Sam: And one last thing before you go: if there's a claim in here you'd push back on, or a thread you want us to pull harder next time, tell us — podcast@connectiveshift.com. We read every single message, and it genuinely decides what we dig into next.

Alex: We'll see you in the next one.