Claude Fable 5.1 and Opus 5 Lead AA-Briefcase
Artificial Analysis shares its latest private benchmark results for frontier AI models.
TLDR
Artificial Analysis announced that Anthropic’s Claude Fable 5.1 and Opus 5 lead its AA-Briefcase evaluation, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra gained roughly 85 Elo points over the prior GPT-5.6 Sol. AA-Briefcase is the company’s private held-out test set for frontier models on realistic agentic knowledge work. It evaluates performance across multi-week projects built by industry experts, each containing many linked tasks and thousands of input files. Grading combines rubric scoring and pairwise comparisons to measure verifiable task success, analytical quality, and presentation quality.
Combined views
23.7K
1 Source, first seen 26d ago