GPT-6 Astra leads overall index, but Claude Fable 5.1 still ahead in software engineering

Epoch AI updated its capabilities index (ECI) yesterday. GPT-6 Astra takes the overall lead with a score of 166 — down from the 169 it received at launch, when only the MirrorCode benchmark had been run in the software-engineering category. Three additional software-engineering benchmarks have since been incorporated, pulling the aggregate score down three points.
In the software-engineering sub-index, Anthropic's Claude Fable 5.1 holds a SWE-ECI of 167 against Astra's 164. The 90% CI are wide, however: 162–179 for Fable and 160–170 for Astra, so the gap is not statistically significant. Each domain score rests on only three to six benchmarks, leaving substantial uncertainty.
Astra's clear advantage appears in Math-ECI, where it sets a new record. Like its software-engineering counterpart, the math sub-index is computed separately from the overall score and draws exclusively on mathematical benchmarks, reflecting a specific capability rather than an average.
For comparison, GPT-5.6 Sol scores 162 on the overall ECI, and Kimi K3 rounds out the top four at 158. The data are available through Epoch's Domain-specific ECI Explorer, which lets users view math and engineering scores independently or build custom variations.
Epoch's methodology separates domain-specific capabilities from the aggregate figure and notes that every sub-index rests on few data points. When confidence intervals overlap, declaring a "leader" in software engineering is more a matter of table placement than a proven difference.