Coding-agent frustrations, from broken code to ignored instructions
Arena says its analysis of 20,840 traces across 29 models found nonworking code in 69.1% of traces containing complaints. Newer models generally drew a more positive balance of feedback.
TLDR
Arena analyzed 20,840 traces from Agent Arena: Code between August 1 and September 13. It reports that among traces containing complaints, 69.1% flagged code not working, 27.1% incomplete output, 27.1% weak finish and usability, and 20.3% ignored instructions. A trace could receive multiple complaint labels. Arena says newer models tend to receive more positive feedback, but models differ in how they disappoint. Fable 5.1 drew fewer “slop” and design complaints than Astra (3.9% versus 5.7% of sampled traces), and fewer complaints about being “flaky” (3.1% versus 5.0%).
Coding-agent frustrations, from broken code to ignored instructions
Arena says its analysis of 20,840 traces across 29 models found nonworking code in 69.1% of traces containing complaints. Newer models generally drew a more positive balance of feedback.