Will it run?
Models

Epoch AI launches benchmark reviews: 9 of 15 fail

By Rae Whitlock Clawpit staff
Epoch AI launches benchmark reviews: 9 of 15 fail

Epoch AI yesterday launched Benchmark Reviews, an initiative to audit the quality of the benchmarks the industry uses to measure language-model capabilities. In the first wave the organization examined 15 benchmarks: 4 earned Verified status, 9 were rated Flawed, and 2 fell into Not Enough Info because insufficient material was available for review.

Three categories, one clear threshold

The classification rests on an internal rubric that defines three verdicts. Verified means the benchmark generally reflects what it claims to measure and any remaining errors do not materially change the results. Flawed signals at least one fundamental defect; in most cases more than 20% of the tasks contain errors that affect measurement accuracy. Not Enough Info is assigned when Epoch could not obtain the required materials; for private benchmarks it will attempt to coordinate a review with the creators while keeping the questions out of the public domain.

The 20% red line

The most striking figure is the 20% threshold: when more than one-fifth of a benchmark’s tasks suffer from errors that influence the score, the benchmark automatically receives a Flawed label. The quantitative bar is intended to replace subjective judgment and reveals how widespread quality problems are in tools that have become the industry’s standard yardstick. Every benchmark marked Flawed will receive a condensed review detailing the flaws found; Verified benchmarks will receive a full review that also covers weaknesses and limitations the team identified.

Conflict of interest and versioning

To avoid conflicts of interest, Epoch does not review benchmarks it created itself and invites external parties to do so. Each review addresses a specific version of a benchmark; if developers release a corrected version, Epoch may re-review it. Findings are shared with creators before publication, and if they request a response, a link to that response will be published alongside the review.

Ongoing program

The effort is not a one-off. Epoch plans to continue publishing reviews on a regular basis, prioritizing benchmarks with high impact and wide adoption as well as those the team considers high-quality but that do not receive sufficient attention. In other words, the project aims to serve as an external control layer on the infrastructure that underpins the leaderboards everyone cites.