GPT-6 Astra and Claude Opus 5.5 reportedly lead Fable 5.1 by about 20 points on Terminal-Bench-Science
Artificial Analysis used the same mini-swe-agent testing setup across all models. It says Qwen3.8 Max (0902), scoring 12%, was the best model it had tested outside OpenAI and Anthropic.
TLDR
In results shared on September 24, 2026, Artificial Analysis reports a roughly 20-point lead for GPT-6 Astra and Claude Opus 5.5 over Fable 5.1 on Terminal-Bench-Science. Its highest-scoring model outside OpenAI and Anthropic was Qwen3.8 Max (0902) at 12%. Comparing maximum-effort runs, it says GPT-6 Sol and Opus 5.5 improved performance while reducing cost per task versus GPT-5.6 Sol and Opus 5, respectively. Artificial Analysis says it used the mini-swe-agent harness across all models to keep comparisons like-for-like.
Combined views
5.8K
1 Source, first seen 16h ago
GPT-6 Astra and Claude Opus 5.5 reportedly lead Fable 5.1 by about 20 points on Terminal-Bench-Science
Artificial Analysis used the same mini-swe-agent testing setup across all models. It says Qwen3.8 Max (0902), scoring 12%, was the best model it had tested outside OpenAI and Anthropic.
TLDR
In results shared on September 24, 2026, Artificial Analysis reports a roughly 20-point lead for GPT-6 Astra and Claude Opus 5.5 over Fable 5.1 on Terminal-Bench-Science. Its highest-scoring model outside OpenAI and Anthropic was Qwen3.8 Max (0902) at 12%. Comparing maximum-effort runs, it says GPT-6 Sol and Opus 5.5 improved performance while reducing cost per task versus GPT-5.6 Sol and Opus 5, respectively. Artificial Analysis says it used the mini-swe-agent harness across all models to keep comparisons like-for-like.