Will it run?
Models

Anthropic cuts a third of cost and pushes fable 5.1 to 90% on arc-agi-2

By Rae Whitlock Clawpit staff
Anthropic cuts a third of cost and pushes fable 5.1 to 90% on arc-agi-2

the result speaks for itself

Anthropic’s Fable 5.1 model achieved 90.0% on the ARC-AGI-2 benchmark and 97.5% on ARC-AGI-1, according to an official ARC Prize post on X. Across the two tests the average cost per task fell by about 32% compared with Fable 5, reaching $3.12 (≈11.5 shekel) per task on ARC-AGI-2 and $1.40 (≈5.1 shekel) on ARC-AGI-1. The improvement is attributed to better token efficiency rather than an increase in model size.

what these benchmarks test

ARC-AGI (Abstraction and Reasoning Corpus) measures visual generalization and inference ability without prior training on the specific tasks. The first version was released in 2019, the second in 2024 with harder problems, and the third is under development. A 90% score on ARC-AGI-2 places Fable 5.1 near the “average human performance” threshold set by the competition organizers at roughly 92% based on published data.

where the money goes

The cost gap between the two benchmarks reflects task complexity: ARC-AGI-2 requires more inference steps and a substantially larger token count per solution. The 32% reduction in average cost comes entirely from cutting the number of tokens the model generates before reaching a final answer, not from any change in Anthropic’s API pricing. In other words, the model “thinks” shorter without losing accuracy.

the obstacle for arc-agi-3

ARC Prize reports that API requests for ARC-AGI-3 have repeatedly been flagged by Anthropic’s systems as reverse-engineering attempts, preventing the test from being completed before publication. Scores for the third version will be released once the evaluation is finished. This is an operational security filter, not a model failure, but it leaves a data gap for now.

what it means going forward

Combining high accuracy with lower cost points to a trend where reasoning models become economically viable for real-time problem-solving applications. Performance metrics for ARC-AGI-3, latency figures, and behavior at the lower end of the task distribution remain unavailable. When the ARC-AGI-3 evaluation is completed, it will be possible to assess whether the efficiency gains persist into the next benchmark tier.