• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Artificial Analysis Announces Intelligence Index v4.2

    AI benchmarking company accelerates v5 elements via interim index updates with new agent tasks.

    AW
    JR
    EM
    17 Sources, 26d ago, first seen 26d ago

    TLDR

    Artificial Analysis posted on X that it released Intelligence Index v4.2. The company said it is accelerating parts of its planned v5 release through interim updates to match frontier models. The post states that v4.2 adds more complex and realistic tasks plus additional private test sets intended to reduce gaming. The changelog entry notes the addition of AA-Briefcase, described as the company's agent. The announcement includes an attached screenshot showing two characters. No further details on scoring changes or model rankings appear in the post.

    Combined views

    1.6M

    17 Sources, first seen 26d ago

    Combined views

    1.6M

    17 Sources, first seen 26d ago

    5.9K likes
    5.9K likes
    471 comments
    733 saves
    561 reposts

    Sentiment

    Positive47.4%52.6%Negative

    Based on 123 sentiment-bearing replies from 116 accounts across 6 conversations.

    471 comments
    733 saves
    561 reposts

    Sentiment

    Positive47.4%52.6%Negative

    Based on 123 sentiment-bearing replies from 116 accounts across 6 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    17 Sources

    @ArtificialAnlysAnnouncing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index v4.2 changelog: + AA-Briefcase, our agentic knowledge work evaluation with a private test set + @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! Intelligence Index v4.2 changes in detail: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5. ➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure. Key results: ➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google ➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier ➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
    @alexandr_wangupdated artificial analysis index—muse spark 1.3 max still performs quite well! the efficient frontier is all Muse, Claude, and GPT
    @EdwardSun0909muse spark 1.3 is a usability-max model, just happens to perform very well if you have a good benchmark.
    @ElaineYaLe6🥑Muse Spark 1.3 max is a strong model! Close to frontier performance, while sitting on the efficiency frontier. We put a lot of care into post-training it. A good reminder that strong fundamentals, solid execution and attention to detail can take you pretty far. Give it a try — feedback is very welcome!
    @echen@ArtificialAnlys just added GDP.pdf to their Intelligence Index. Which means their definition of intelligence now includes something deceptively simple: can models understand the documents you deal with on a normal Tuesday at work? Leases, invoices, dosage tables, and financial reports. Astra, the best model, still solves just under 1 in 3. Master the boring, master the frontier.
    @firstadopterAA updates their flagship index: "Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming" Gemini still at number 10
    @emollickThe whole idea of indexes that you don't change all the criteria in ways that hugely change the rankings and evaluations of existing models. (Also GDPval-AA remains a terrible measure and AA should have removed it when they developed their own Briefcase-AA benchmark)
    @_weipingFirst time in my life seeing a benchmark scramble overnight to fit a model. It used to be benchmaxxing; now it’s model-maxxing 🤣 Build a great model, and the leaderboard will chase after you 😉
    @MLStreetTalkRT @ArtificialAnlys: Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with i…
    @jack_w_raeRT @alexandr_wang: updated artificial analysis index—muse spark 1.3 max still performs quite well! the efficient frontier is all Muse, Cla…

    17 Sources

    @ArtificialAnlysAnnouncing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index v4.2 changelog: + AA-Briefcase, our agentic knowledge work evaluation with a private test set + @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! Intelligence Index v4.2 changes in detail: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5. ➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure. Key results: ➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google ➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier ➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
    @alexandr_wangupdated artificial analysis index—muse spark 1.3 max still performs quite well! the efficient frontier is all Muse, Claude, and GPT
    @EdwardSun0909muse spark 1.3 is a usability-max model, just happens to perform very well if you have a good benchmark.
    @ElaineYaLe6🥑Muse Spark 1.3 max is a strong model! Close to frontier performance, while sitting on the efficiency frontier. We put a lot of care into post-training it. A good reminder that strong fundamentals, solid execution and attention to detail can take you pretty far. Give it a try — feedback is very welcome!
    @echen@ArtificialAnlys just added GDP.pdf to their Intelligence Index. Which means their definition of intelligence now includes something deceptively simple: can models understand the documents you deal with on a normal Tuesday at work? Leases, invoices, dosage tables, and financial reports. Astra, the best model, still solves just under 1 in 3. Master the boring, master the frontier.
    @firstadopterAA updates their flagship index: "Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming" Gemini still at number 10
    @emollickThe whole idea of indexes that you don't change all the criteria in ways that hugely change the rankings and evaluations of existing models. (Also GDPval-AA remains a terrible measure and AA should have removed it when they developed their own Briefcase-AA benchmark)
    @_weipingFirst time in my life seeing a benchmark scramble overnight to fit a model. It used to be benchmaxxing; now it’s model-maxxing 🤣 Build a great model, and the leaderboard will chase after you 😉
    @MLStreetTalkRT @ArtificialAnlys: Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with i…
    @jack_w_raeRT @alexandr_wang: updated artificial analysis index—muse spark 1.3 max still performs quite well! the efficient frontier is all Muse, Cla…