Introducing the Artificial Analysis Intelligence Index v4.2

0
10

Key Takeaways

  • The Intelligence Index v4.2 is an interim update designed to keep the benchmark aligned with rapid frontier‑model progress while preventing evaluation gaming.
  • Two new evaluations—AA‑Briefcase (agentic knowledge‑work) and GDP.pdf (long‑context document reasoning)—introduce more realistic, multi‑step tasks and private held‑out test sets.
  • Held‑out data now accounts for 40 % of the Index weighting, double the previous share, and will grow further in v5.
  • Grading infrastructure has been upgraded (system prompts, corrected answer keys, re‑anchored Elo scales, sandbox robustness) to improve scoring stability and accuracy.
  • Anthropic’s Claude Fable 5.1 leads the overall Index, closely followed by OpenAI’s GPT‑6 Astra; Meta, SpaceXAI, Moonshot/Kimi, Z AI, and Google round out the top labs.
  • Four labs—Anthropic, OpenAI, Meta, and Z AI—share the updated Cost‑per‑Task Pareto frontier, indicating comparable efficiency at the intelligence frontier.
  • GPT‑6 Astra dominates the Output‑Token Pareto frontier, being far more token‑efficient than peers near the top of the Index.
  • The team confirms ongoing work on Index v5 and promises additional incremental releases in the near future.

Overview of the Interim Update
The Intelligence Index has received an interim bump to v4.2 to “keep pace with the frontier” as model capabilities evolve rapidly. As the article states, “We have been planning and building elements of Index v5 for months – it’s been 8 months since we launched Index v4 in January.” The update arrives because recent major model launches have moved the frontier quickly, and the team felt it important to deliver an immediate update so the Index “remains as relevant and useful as ever to users.” By holding back larger changes until v5, the developers aim to preserve stability while still reflecting the latest advances.


Introducing AA‑Briefcase: Agentic Knowledge‑Work Evaluation
A centerpiece of v4.2 is the addition of AA‑Briefcase, an in‑house evaluation that mimics real‑world agentic knowledge work. The description notes that AA‑Briefcase “tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi‑week knowledge work projects, each with many linked tasks and thousands of input source files.” It combines rubric‑based and pairwise grading to assess verifiable task success, analytical quality, and presentation quality, yielding a holistic view of agentic capability. As the source explains, “AA‑Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.”


GDP.pdf: Long‑Context Document Reasoning
Complementing AA‑Briefcase, the GDP.pdf evaluation—created by Surge AI—challenges models to reason across extremely long documents. The article details that GDP.pdf “evaluates single‑turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions.” Responses are judged against 1,275 expert‑authored atomic criteria, with the headline All‑pass Rate awarding credit only when every criterion is satisfied. This design pushes models to handle extensive, heterogeneous information in a single turn, reflecting tasks such as legal briefs, technical reports, or financial analyses.


Weighting Shift to Prevent Gaming
To curb the tendency of labs to over‑fit to public test sets, v4.2 increases the proportion of private, held‑out data. The update specifies that “40 % of our Index weighting is now private, held‑out test sets – double the figure from v4.1.” Held‑out components include AA‑Briefcase, AA‑Omniscience, and solutions for CritPt. By expanding the hidden share, the Index reduces the ability for labs to “game evaluations,” and the article notes that “the held‑out percentage will increase further in Index v5.” This shift aims to make scores more reflective of genuine generalization rather than memorization.


Grading Infrastructure Enhancements
Reliability of scoring received concrete upgrades. In AA‑LCR v1.1, the team “added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy.” For GDPval‑AA v2 and AA‑Briefcase, they “improved our sampling and re‑anchored the Elo scale, making ratings more stable as new models are added.” Additionally, for SciCode, grading sandboxes were hardened “to ensure slow but correct code does not count as a failure.” These changes collectively aim to produce more consistent, trustworthy ratings as the model landscape expands.


Leaderboard Highlights: Anthropic and OpenAI Lead
The updated Index places Anthropic’s Claude Fable 5.1 at the summit, followed closely by OpenAI’s GPT‑6 Astra. The article quotes, “Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT‑6 Astra, which shows a 4pt gain over GPT‑5.6 Sol.” Meta ranks third, with SpaceXAI, Moonshot/Kimi, Z AI, and Google completing the top tier. In the AA‑Briefcase leaderboard, Claude Fable 5.1 and Opus 5 lead, while GPT‑6 Astra outpaces GPT‑5.6 Sol by roughly “85 Elo points.” On GDP.pdf, OpenAI’s GPT‑6 Astra scores 33.2 % All‑pass, edging out GPT‑5.6 Sol (28.2 %) and Claude Fable 5.1 (26.2 %).


Cost‑per‑Task Pareto Frontier
Efficiency considerations are captured in the Cost‑per‑Task Pareto frontier, which now features four labs at its edge. The article notes, “Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier.” This indicates that, among the leading contenders, these organizations achieve comparable trade‑offs between raw intelligence scores and the computational or financial cost required to attain them. The frontier will likely shift as v5 introduces new weighting and tasks, but for v4.2 the balance is shared across these four players.


Output‑Token Pareto Frontier: GPT‑6 Astra’s Token Efficiency
When measuring intelligence per output token, GPT‑6 Astra stands out. The source states, “GPT‑6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash‑Lite at either end of the curve (excludes models below 25 on the Index).” A later clarification adds that “GPT‑6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1 and Gemini 3.8 Flash using the most tokens of models scoring at least 25 on our Index.” This suggests that, while some models may achieve high scores, they do so at a substantially higher token cost, making GPT‑6 Astra the most economical choice for applications where output length matters.


Looking Ahead: Toward Index v5
Despite the interim nature of v4.2, the team emphasizes that work on the next major version is well underway. The article concludes, “Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned!” This roadmap signals that users can expect further refinements—likely greater reliance on held‑out sets, additional real‑world task simulations, and continued upgrades to grading robustness—as the benchmark strives to stay aligned with the ever‑accelerating frontier of AI capabilities.


All quoted passages are taken directly from the source article provided.

https://www.linkedin.com/pulse/announcing-artificial-analysis-intelligence-index-jsvnc

SignUpSignUp form

LEAVE A REPLY

Please enter your comment!
Please enter your name here