Will it run?
Models

Artificial Analysis adds GPT-6 Sol and Luna to independent benchmark

By Rae Whitlock Clawpit staff
Artificial Analysis adds GPT-6 Sol and Luna to independent benchmark

The independent evaluation platform Artificial Analysis has updated its leaderboards to include GPT-6 Sol and Luna, the two newest models from OpenAI, alongside the market's leading competitors. The site, which focuses on independent testing rather than relying on vendor-supplied data, now offers a more comprehensive snapshot for anyone choosing a model and provider for production.

The comparison centers on the Artificial Analysis Intelligence Index, which ranks models by performance across diverse benchmarks and cross-references those scores against the average cost per task on the index. The chart makes it immediately clear which model delivers more "intelligence per dollar" and separates open-weights models from proprietary ones. A historical tracker alongside it shows how the model frontier has advanced over time.

Two relatively new agentic benchmarks now feature prominently. AA-Briefcase v1.1 tests agents on long-horizon knowledge tasks — creating spreadsheets, presentations and realistic business memos. AA-AnalystAgent measures end-to-end quantitative analysis performance on real spreadsheets and documents of the kind business and data analysts handle daily. Both benchmarks report scores in Elo and pass^5 respectively, enabling stable statistical comparison.

AA-Omniscience targets the hallucination problem: the benchmark rewards accuracy, penalizes bad guesses, and provides a broad view of factual reliability across domains. GDPval-AA v2.1 takes a different angle, evaluating models on tasks with real economic value across a wide range of professions. Both metrics are being updated to new versions — v2.1 for GDPval, v1.1 for Briefcase — signaling rapid iteration in testing methodology.

The platform's Openness Index ranks models by the availability and transparency of various components — weights, code, training data and licenses — and cross-references that against the Intelligence Index. In parallel it publishes speed data (latency, time to first token), actual token throughput on an intelligence task, and full pricing breakdowns: input price, output price and cache-hit price. All measurements are taken through the providers' official APIs, not through chat interfaces.

For developers and companies selecting a model for production, the update provides a broader comparison basis than before: not just an overall score but a breakdown by task type (coding, analysis, business writing, factual knowledge), by marginal cost, by response speed and by degree of openness. GPT-6 Sol and Luna now enter that test suite, and their initial results will appear on the leaderboards in the coming days.