Artificial Analysis launches unified coding-agent index: three benchmarks, one score

Artificial Analysis has published the Coding Agent Index, a weighted score that merges three public benchmarks into a single view of code-agent performance. The index combines DeepSWE v1.1, with 113 software-engineering tasks; Terminal-Bench 4.0, with 66 terminal-work tasks; and SWE-Atlas-QnA, with 124 technical questions about code repositories. Each component probes a different capability: architecture understanding and code reading; producing fixes and patches that pass tests; and multi-step command-line environment navigation.
Methodology: normalized pass@1 per task
The calculation uses a task-level normalized pass@1 average. For each task the agent runs three attempts; the task score is the mean of those attempts, and the final score is the mean across all tasks so every task carries equal weight. All outcomes are binary — success or failure — and an attempt can finish "cleanly" yet still score zero if it fails the verifier. In SWE-Atlas-QnA the method aligns with Scale AI's Task Resolve Rate. This approach avoids bias toward unusually easy or hard tasks, but it also discards nuance about partial-fix quality.
Cost, tokens, and runtime, not just score
Alongside the composite score Artificial Analysis publishes four additional axes: cost per task at current API pricing (including cache discounts), total token consumption broken down by type, active runtime per task, and safety-refusal rate. The data let users compare efficiency — how many tokens and dollars are burned to solve a task — not just who finishes first. Many users will access agents through subscriptions rather than per-token billing, so the shown cost is an indicator, not a final invoice.
Reward hacking: when the agent games the benchmark
A newer metric on the page is Reward Hacking Rate, the frequency with which an agent receives credit for a task without demonstrating the capability the task measures. This is a known reinforcement-learning problem: the model learns to exploit verifier loopholes instead of solving the problem. Publishing the rate next to the main score forces a fairer comparison, because a high score with a high hacking rate is worth less than a lower score with zero exploitation.
Harness comparison, still not here
The platform promises a comparison of the different harnesses that run the agents (coming soon). That is critical because the same model on a different harness can yield dramatically different results — context management, execution permissions, and editing tools change the game. Until that analysis appears, the current index reflects performance in the specific configuration tested, not an intrinsic property of the model alone.
What's missing: no raw performance numbers in the table
The page shows charts and trends but no table with the raw numbers for each agent on each benchmark separately. Without that, it is hard to reconstruct the calculation or check whether an agent is strong on DeepSWE but weak on Terminal-Bench. For users who need a specific capability — say, bug fixes without Q&A — the composite index may hide the real fit.