Claude Opus 5 hits 100% on accounting benchmark; human CPAs average 37%

Claude Opus 5 scored a clean 20 out of 20 on a set of well-defined accounting tasks, finishing each one in minutes. The human comparison group — 12 licensed CPAs — ranged from 0% to roughly 90%, with a mean of 37%. About two years ago, the same tasks yielded near-zero results from GPT-4o (mid-2024 vintage).
The jump happened inside 18 months. According to a post on the Mercor blog, frontier models still trailed human accountants by a wide margin a year and a half ago. The swing from near-zero to 100% compressed into an unusually short window, and the model overtook the human average sometime in 2025. The researchers put it bluntly: "In medium-length, well-defined accounting tasks, frontier models are faster and more accurate than junior CPAs, even the best one in our study."
Caveats are baked into the methodology. The test covered rule-based tasks of medium length — the bread-and-butter of junior accounting work. The human sample is small (12 people), and the tasks were deliberately well-specified, which plays to LLM strengths. The authors did not publish numbers for open-ended, long-horizon work or tasks that demand broader professional judgment.
The takeaway is straightforward: in roles where the work is governed by clear rules, humans have become the slower, costlier, more error-prone option. Opus 5 cleared every task in minutes; the CPAs needed significantly more time and showed far wider variance in accuracy. The result doesn't prove the profession is replaceable end-to-end, but it draws a sharp line — the algorithmic slice of the job no longer needs a human operator.