Sonnet 5.5 (max) reportedly led a two-repo bug test, but used far more turns
A user testing models against 105 bugs across two real repositories gave Sonnet 5.5 (max) a score of 55.5, ahead of GPT-6 Astra (max) at 45. The user listed 1,330 turns for Sonnet versus 222 for Astra.
TLDR
In one user's test involving 105 bugs across two repositories, Sonnet 5.5 (max) scored 55.5, compared with 45 for GPT-6 Astra (max) and 41.7 for Opus 5.5 (max). The user listed 1,330 turns for Sonnet across two runs, 222 for Astra across three and 476 for Opus across three. Sonnet's xhigh setting scored 39 with 588 turns in one run.
Combined views
133
1 Source, first seen 1h ago
