Will it run?
Models

Artificial Analysis publishes full methodology for Intelligence Index v4.2

By Rae Whitlock Clawpit staff

Artificial Analysis has released the complete methodology behind version 4.2 of its Intelligence Index, an evaluation suite that aggregates ten separate benchmarks into a single score. The index covers reasoning, knowledge, mathematics and coding, weighted across four categories: agents at 30%, coding at 20%, scientific reasoning at 20% and general at 30%. The emphasis on agentic tasks reflects where the market is heading — fewer trivia questions, more multi-step execution.

The benchmark lineup comprises AA-Briefcase, GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR v1.1, AA-Omniscience, Humanity's Last Exam, GDP.pdf and CritPt. Every model is tested under identical conditions: same prompts, temperature 0 for standard models and 0.6 for reasoning models, and a 16,384-token output cap for standard models or the lab-declared maximum for reasoning models. When an API call fails, the system retries up to 30 times before abandoning the question.

The team says the 95% confidence interval for the overall index is under ±1%, based on more than ten repeat runs per model on each benchmark. That is unusually tight for the field, though the company cautions that individual benchmark results may show wider intervals. A more detailed statistical write-up is expected later.

Multilingual ability is excluded from the main score. It lives in a separate Artificial Analysis Multilingual Index, built on Global-MMLU-Lite and spanning 16 languages: English, Chinese, Hindi, Spanish, French, Arabic, Bangla, Portuguese, Indonesian, Japanese, Swahili, German, Korean, Italian, Yoruba and Burmese. Vision and speech capabilities are also measured separately and not folded into the core index.

The methodology rests on four principles: full standardization of test conditions, no unfair penalty for answers that correctly follow instructions, zero-shot evaluation only — no examples in the prompt — and transparency, including publication of prompt templates, evaluation criteria and limitations. That approach aligns with how modern chat models are designed to work: clear instructions, no few-shot.

Practically, this means there is now a systematic, well-documented baseline for comparing models — something that has been missing when you only look at individual benchmark tables. The heavy weighting on agency and coding pushes vendors to optimize there, while keeping multilingual and visual skills separate prevents conflating distinct capabilities in one number. The open question is how well this score correlates with real-world production utility, and that is a test only time and developers will answer.