Will it run?
Models

DeepSeek redraws the efficiency frontier in Agent Arena

By Rae Whitlock Clawpit staff
DeepSeek redraws the efficiency frontier in Agent Arena

DeepSeek released V4.1-Flash (Max) and it entered directly into third place among open models in Agent Arena, the benchmark that simulates multi-step agent tasks. The real achievement is not the raw score — a 4.87% net improvement over the arena baseline — but the median cost per task: 7 cents. At that point the model redefines the Pareto frontier, the line separating models that deliver better value per dollar from those that do not.

The numbers behind the ranking

The full leaderboard shows Claude 5.1 (Max) at 13.9% improvement for $4.54 per task, GPT-6 Astra (Max) at 11.9% for $4.09, and Claude Opus 5 (Max) at 11.09% for $3.52. DeepSeek's model delivers a little more than one-third of the leaders' improvement at roughly one-sixtieth the price. Even against Claude Opus 4.8 (Hi) at $1.36 per task, the gap is 22x in cost favoring the Chinese model.

Comparison within the open tier

Inside the open-model group the contrast is sharper. Hy4 Preview scores 4.96% improvement at 22 cents per task; DeepSeek retains 98% of that performance at 73% lower cost. Kimi K3 (Max) reaches 6.39% at 77 cents, and DeepSeek holds 76% of the performance at a 92% discount. In net cost-benefit terms, V4.1-Flash (Max) is currently the most efficient model in the top three.

Models pushed off the frontier

DeepSeek's entry pushed three models off the Pareto frontier: GPT-5.6 Luna (Extreme-Hi), GLM-5.3-Flash, and the previous generation DeepSeek-V4-Flash. The practical implication: for each of them there is now another model that delivers equal or better performance at a lower price, or significantly better performance at the same price. In an arena where developers pick models against a hard budget, leaving the frontier means immediate loss of relevance.

What this means in practice

For anyone building autonomous agents that run long loops — code scanning, task planning, repeated API calls — the gap between 7 cents and $4 per task is the difference between a pilot shut down after a week and a product that ships to production. DeepSeek does not claim the absolute performance crown; it claims the ratio crown, and the data currently back the claim.