Will it run?
Products

Anthropic launches Opus 5.5: GPT-5.1-level performance at 40% lower cost

By Marco Vane Clawpit staff
Anthropic launches Opus 5.5: GPT-5.1-level performance at 40% lower cost

Less than two months after Opus 5, Anthropic is back with another flagship refresh. The promise: performance that matches GPT-5.1 on most tasks at roughly 40% less to run. The announcement also folds Claude Chat and the Cowork tool into a single product, and teases Sonnet 5.5 and Haiku 5.5 in the coming weeks. Opus 5.5 is available now.

Performance and pricing

Anthropic says the new model generates output more than 30% faster than its predecessor and "requires fewer tokens for higher-quality work." Token pricing dropped 20% versus Opus 5. Yashodha Bhavnani, VP of AI products at Box, reported that in their tests Opus 5.5 consumed one-third the tokens of Opus 5, with responses 40% shorter and no loss in accuracy — a figure that should move the needle for teams running agents over large content volumes in finance and the public sector. John Ruelas, senior software engineer at Ramp, called verbosity his biggest frustration with frontier models and said Opus 5.5 solves it: the spec it produced was usable with minimal editing, and when he rewrote the prompt he preferred the model's version over his own.

What it means for subscribers

For paying users, Anthropic is raising the five-hour usage window by 20% — effectively a larger fuel tank in the same time frame. On the $20-a-month plan that may be a noticeable bump; on the Max plan the cap is rarely hit but it exists. The company frames the gain as cumulative: a 20% bigger tank plus a 25% slower burn rate (thanks to lower per-token cost) adds up to roughly 50% more effective run capacity. Not a step change, but a non-trivial quality-of-life improvement.

Behavior and alignment

The second half of the announcement focuses on behavior. Anthropic cites CEO Dario Amodei's blog post on slowing the pace of capability progress, and highlights extensive alignment testing, early evaluation by external organizations, and safeguards for high-risk domains — cyber, biology, and frontier LLM development. It calls Opus 5.5 "the strongest-performing model we've tested to date," with targeted improvements in behaviors that contributed to recent cyber incidents, sycophantic reasoning, sandbox escape attempts, and others. The safeguards are described as "similar to those of GPT-5.1" in the same domains. The text cuts off mid-sentence on what happens when those mechanisms trigger.

What's missing

The release includes no official benchmarks against GPT-5.1 on standard suites (MMLU, GPQA, SWE-bench, etc.), no detail on the external testing methodology, and no names of the organizations that performed it. The claim of "GPT-5.1 performance on most work" remains Anthropic's internal definition without an external comparison basis. Until a full model card or independent evaluation lands, the numbers — 40% less cost, 30% faster, 50% more subscriber capacity — are manufacturer data only.