Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5
Sep 22, 2026, 11:59 AM · MarkTechPost
Anthropic’s Claude Opus 5.5 opens the 5.5 family claiming Fable 5.1-level work at roughly 40% lower running cost than Opus 5 — with tighter safeguards and thinking that won’t switch off.
Why it matters
MarkTechPost’s September 22, 2026 report: Claude Opus 5.5 is the first model in Anthropic’s Claude 5.5 family, positioned at Claude Fable 5.1 level on most work and about 40% cheaper to run than Opus 5 on typical workloads at default settings. API id claude-opus-5-5 on Claude Platform, AWS, Google Cloud, and Azure; closed weights; zero data retention available.
Pricing versus Opus 5: input $4 vs $5, output $20 vs $25, cache reads $0.20 vs $0.50 (60% drop), cache writes $5 vs $6.25. Anthropic says fewer tokens per task plus cheaper cache reads drive the ~40% workload savings; output generation more than 30% faster. Fast mode up to 2.5× speed at $8 / $40.
From the desk
We’re reading Opus 5.5 as Anthropic’s efficiency counterpunch in the same launch window as OpenAI’s Sol/Luna — and as the first major drop since CEO Dario Amodei publicly pushed for pacing the frontier.
On Anthropic-reported benches with adaptive thinking at max effort and safeguards on, Opus 5.5 leads several agentic coding and knowledge-work scores: Terminal-Bench 4.0 at 66.4%, FrontierCode v1.1 54.4%, CursorBench 4.0 57.8%, GDPval-AA Elo 1846, OSWorld 2.0 81.8%. GPT-6 Astra still leads Terminal-Bench-Science (64.6% vs 58.7%) and AutomationBench (41.4% vs 40.0%). Anthropic itself warns benchmark margins are getting less reliable and that the practical gap to Fable 5.1 is narrower than the table. Cost-adjusted framing is sharper: at medium effort, FrontierCode 54.6% beats Astra’s 53.3% top at about a fifth the cost per task.
Early tester anecdotes are aggressive — 680k-line migration in under a day; 200k-line audit/fix under three hours versus Opus 5’s 20+ hours and 2.5× tokens; HAProxy C-to-Rust port in 9.5 hours at 51% less cost than Fable 5.1; Deloitte lowest-effort Opus 5.5 catching 72% of known review bugs versus 56% for Opus 5 at high effort. Useful AI wants exactly this: stronger work per dollar, faster loops, clearer writing that puts key information first.
Eyes open on the safeguard stack. Biology and cyber capabilities are framed as comparable to Claude Mythos 5.1, so Fable-class controls apply: most non-routine cybersecurity tasks re-route to Opus 4.8 pending Cyber Verification Program expansion; Life Sciences Verification for vetted orgs; preserved thinking blocks API users (accounts on/after August 31, 2026) from editing prior context to extract reasoning; thinking cannot be disabled; EU AI Act watermarking on outputs. Best score yet on Anthropic’s ~2,000-scenario behavioral audit; ~85% fewer containment circumvention attempts than Opus 5 in a new test — and the model often suspects it’s being evaluated.
I’m watching independent cost-per-task replications, whether cyber/bio routing frustrates legitimate security researchers, and whether “thinking always on” plus watermarking become the default enterprise expectation after this release.
Context
Asif Razzaq / MarkTechPost. External pre-release testing noted from METR and Frontier Design. Anthropic raising five-hour usage limits on Pro, Max, Team, and seat-based Enterprise, with a saveable rate-limit reset. Full details pointed to the Opus 5.5 System Card.
Who feels it
- Coding agent and IDE teams
- CursorBench / Terminal-Bench / FrontierCode claims plus 40% cost cut force an immediate re-bake of default model choices.
- Enterprise security and compliance
- Always-on thinking, watermarking, preserved thinking, and cyber/bio re-routes change integration assumptions.
- Cloud marketplace buyers
- Same model on Claude Platform, AWS, GCP, and Azure — pick the compliance boundary, not the capability.
- Safety community
- First major Anthropic drop after Amodei’s pacing call — containment and audit numbers will be parsed as signal of whether slowdown talk matches shipping practice.
What to watch
- Third-party AutomationBench / Terminal-Bench / OSWorld replications versus Astra and Fable 5.1.
- Cyber Verification Program expansion timeline for Opus 5.5.
- Real workload cost deltas once cache-read mix is measured outside Anthropic’s “typical” framing.
- Developer reaction to non-disableable thinking and output watermarking.
Companies: Anthropic