Claude Sonnet 5.5 lifts Anthropic to top of Agent Arena
Anthropic's Claude Sonnet 5.5 (Max) entered the Agent Arena leaderboard directly at number three with a 12.5% net improvement, giving the company a clean sweep of the top three positions. The jump represents an 8.1 percentage-point gap over the previous version, Sonnet 5 (High), which sits at number 13 with only a 4.4% net improvement.
In the chat category, Sonnet 5.5 takes first place with a 15.6% improvement, ahead of Phi-5.1 at 11.49% and Opus 5.5 at 10.29%. The performance comes at a steep price: a median cost of $2.74 per task, 73% higher than the nearest rival, Opus 5.5 (High), at $1.58 per task. Organizations will need to weigh that cost-benefit gap carefully before putting the model into production.
On the coding front, the xHigh-reasoning version scored 1,786 points in Code Arena: WebDev, landing in third place — just two points behind GPT-6 Astra in second. At a blended price of $8 per million tokens, the model remains on the Pareto frontier, meaning no other model currently delivers more performance for less cost, or the same output for less money, with enhanced reasoning enabled.
The dominance across the top three of Agent Arena suggests Anthropic's current architecture is well tuned for agentic tasks — planning, tool use, and self-correction loops. But the cost gap versus Opus 5.5 (High) is a reminder that the performance leap isn't free. Engineering teams will have to decide whether the extra 12.5% justifies nearly doubling their inference budget.