Will it run?
Models

GPT-6 Astra claims unprecedented cost efficiency and record ARC-AGI scores

By Rae Whitlock Clawpit staff
GPT-6 Astra claims unprecedented cost efficiency and record ARC-AGI scores

An X user called cheaty has posted a thread claiming OpenAI's GPT-6 Astra redraws the Pareto frontier of cost efficiency — extreme token efficiency makes it cheaper per task than Gemini 3.8 Flash, a model that costs 13 times less. The source says Astra cuts output tokens by roughly 10% at max effort compared to GPT-5.6 Sol while keeping a higher intelligence index.

The most interesting bit: the cost curve on ARC-AGI bends backward. Max reasoning level solves problems so fast the total cost drops below what you'd pay at lower reasoning levels. If true, extra inference compute pays for itself by dramatically shrinking answer length.

On ARC-AGI-3, the benchmark from François Chollet that tests generalization to entirely novel problems without prior training, Astra reportedly hits 63% standard and 99% with a new "provider adapter harness." The model allegedly beats human performance on 96% of levels and builds the most accurate symbolic model of new environments seen so far.

All of this comes from one social-media account. No OpenAI release, no technical paper, no independent replication. "Pareto frontier" means you can't improve one metric without hurting another — a strong claim that needs detailed numbers. The Gemini 3.8 Flash comparison is also hard to check since full details aren't public. Until open benchmarks or an official report land, the numbers are just unverified claims.