Will it run?
Models

Anthropic hits 1.5x R&D acceleration, METR estimates

By Rae Whitlock Clawpit staff
Anthropic hits 1.5x R&D acceleration, METR estimates

METR, an organization that specializes in evaluating model capabilities and risks, estimates that Anthropic has reached a 1.5x acceleration rate in AI-driven research and development. The definition behind the number is strict: this is not point assistance but a workflow in which most code, experiments, and experimental decisions are written and defined by models, with human contribution shrinking to high-level steering only.

METR's methodology tags every commit with an LLM-assigned confidence score, automatically prioritizes experiments that cross a predicted-impact threshold, and routes low-risk code through a continuous generation pipeline. The result, they claim, is orders of magnitude more model-originated lines of code and experiments than those written by humans from scratch.

Not everyone buys the figure. The researcher Aifp notes that 1.5x roughly matches the AC milestone in the AI Futures model but estimates the real pace is still far from that. Both sides agree the list of caveats is longer than usual: the measurement has not been fully published, definitions vary across labs, and no agreed benchmark yet allows direct cross-company comparison.

Leaked technical details — commit tagging, experiment prioritization, an automatic generation pipeline — point to mature engineering infrastructure. But without a full METR report it is impossible to know how much of the acceleration comes from automating repetitive tasks versus genuine scientific breakthroughs. The expectation now is a formal publication from the "AI R&D Acceleration Measurement Team" that would provide a transparent methodology and comparable data.

If the number holds up to external scrutiny, the practical implication is that Anthropic is compressing a two-month development cycle into three weeks, a jump that changes the economics of large-language-model research. Until then, 1.5x remains an internal estimate with a promising but unverified methodology, and the research community is waiting for data that can be critically examined.