GPT-6 Astra costs more and delivers mixed results against its predecessor

A new Artificial Analysis breakdown shows GPT-6 Astra running 75% more expensive than GPT-5.6 Sol at max effort, and it sits mostly behind the older model on the Intelligence Index efficiency frontier — intelligence per dollar per task. The gap comes from a 2.5× price increase that is only partly offset by lower token consumption. In plain terms: you pay far more for similar or worse per-dollar performance on the general benchmark.
The picture flips on the Coding Agent Index. There Astra posts a significant jump, matching Fable 5’s score at a lower cost. At max effort it lands on the Pareto frontier of coding intelligence versus cost per task: priced roughly the same as GPT-5.6 Sol at max effort but scoring two points higher. Astra also defines a new Pareto frontier on the Intelligence Index versus output tokens per task, cutting output tokens by about 10% compared with its predecessor.
The sharpest improvement appears on the AA-Omniscience benchmark: hallucination rate drops from 92% to 51% at max effort, accompanied by a modest accuracy gain. That steep decline in hallucinations drives the overall score upward and signals a fundamental behavioral shift — fewer fabrications, tighter adherence to facts.
Results on agentic knowledge-work tests are mixed. On AA-Briefcase the model gains roughly 80 points, with significant lifts in both rubric scores and Analytical Quality Elo, yet the Presentation score falls. On GDPval-AA v2 it regresses by a similar magnitude. The takeaway: analysis and reasoning improve, but presentation and execution of complex, long-horizon tasks weaken.
Bottom line: GPT-6 Astra is not a uniform upgrade. It is substantially more expensive, leads on coding and hallucination reduction, but trails on general cost-efficiency and several agency benchmarks. For developers building code agents the premium may be justified. For general knowledge work, GPT-5.6 Sol still delivers more value per shekel.