Artificial Analysis Intelligence Index v4.2: What Changed
The Artificial Analysis Intelligence Index — the ranking procurement teams actually cite — just got its first update in eight months, and the headline isn't a new model. It's 40% private, held-out test weighting, double v4.1, plus the removal of GPQA Diamond, which the index now calls "saturated." The ruler that measures the frontier just got rebuilt to survive it.
What shipped in Index v4.2
Three additions, one removal. AA-Briefcase: an in-house agentic knowledge-work evaluation on a private test set — multi-week projects with thousands of source files, graded by rubric plus pairwise comparison. Surge AI's GDP.pdf: single-turn document reasoning across 100 PDFs and ten domains, evidence scattered over 4,592 pages, graded against 1,275 expert-written atomic criteria — a task only counts when every criterion passes. The removal: GPQA Diamond, once the gold standard for scientific reasoning, dropped as saturated. Behind it: grading infrastructure rework — AA-LCR v1.1, re-anchored Elo in GDPval-AA v2, sandbox fixes in SciCode so slow-but-correct code stops counting as failure.
Why the anti-gaming bar matters
The quiet part, loudly: 40% of every model's score now comes from held-out data labs can't train on or tune against. v4.1 sat at 20%. Artificial Analysis says it held v5 back "to keep the Index stable" — then shipped this because the frontier moved too fast to wait. The message to labs is blunt: your public-benchmark wins are worth half what they were yesterday. The Benchmarkpocalypse isn't cancelled; it just got a private-test firewall.
Who leads the v4.2 leaderboard
Anthropic's Claude Fable 5.1 leads the Index, with OpenAI's GPT-6 Astra showing a 4-point gain over GPT-5.6 Sol. Meta is third, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google. The cost-per-task Pareto frontier is shared by four labs: Anthropic, OpenAI, Meta, and Z.AI. Same mix as before — but now measured against answers nobody has seen.