Artificial Analysis releases its full evaluation map

The independent platform Artificial Analysis (Artificial Analysis) unveils the complete test suite it uses to rank language models, offering a broader picture than standard benchmarks. The site, which focuses on independent comparisons, divides the assessments into four main categories—intelligence, multimodal capabilities, cost and speed—each broken down into specific metrics with 95% confidence intervals.
Intelligence, cost over time and agent testing
At the core of the system is the Artificial Analysis Intelligence Index, which separates models with open weights from proprietary ones and shows the cost per index task. Adjacent is a graph of intelligence development over time and a comparison of coding agents on end-to-end software-engineering tasks, including runtime and cumulative cost. The data allow observers to see not only which model is “smarter” but which delivers intelligence per shekel.
Four benchmarks dedicated to knowledge work and factual accuracy
Four relatively new tests cover concrete business scenarios: AA-Briefcase evaluates agents on long-horizon tasks that require outputs such as spreadsheets, presentations and memos, using an Elo rating; AA-AnalystAgent measures end-to-end quantitative analysis on real spreadsheets and documents with a pass^5 metric; AA-Omniscience focuses on knowledge and hallucinations, rewarding accuracy and penalising poor guesses; and GDPval-AA v2 assesses tasks with real economic value across various professions. Each is designed to reflect daily work rather than academic puzzles.
Openness, token output, pricing and latency
The platform’s Openness Index ranks models by the availability and transparency of components—weights, code, data, license—and lets users compare it with the intelligence index. At the same time, token output per task, detailed cost breakdowns (Cache Hit, input, output) and latency and performance of each provider’s official API are published. The combination gives engineers a complete view: not only how smart a model is, but how much it costs, how fast it runs, and how truly open it is.