• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Claude Fable 5.1 and Opus 5 Lead AA-Briefcase

    Artificial Analysis shares its latest private benchmark results for frontier AI models.

    AA
    1 Source, 26d ago, first seen 26d ago

    TLDR

    Artificial Analysis announced that Anthropic’s Claude Fable 5.1 and Opus 5 lead its AA-Briefcase evaluation, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra gained roughly 85 Elo points over the prior GPT-5.6 Sol. AA-Briefcase is the company’s private held-out test set for frontier models on realistic agentic knowledge work. It evaluates performance across multi-week projects built by industry experts, each containing many linked tasks and thousands of input files. Grading combines rubric scoring and pairwise comparisons to measure verifiable task success, analytical quality, and presentation quality.

    Combined views

    23.7K

    1 Source, first seen 26d ago

    Combined views

    23.7K

    1 Source, first seen 26d ago

    98 likes
    98 likes
    3 comments
    9 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    9 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @ArtificialAnlysAnthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a substantial gain above GPT-5.6 Sol of ~85 Elo points. AA-Briefcase is our frontier in-house evaluation with a private held-out test set. The evaluation tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.

    1 Source

    @ArtificialAnlysAnthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a substantial gain above GPT-5.6 Sol of ~85 Elo points. AA-Briefcase is our frontier in-house evaluation with a private held-out test set. The evaluation tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.