Report
Code understanding may matter more than edit size in coding-agent failures
A post on Microsoft research says reading and analysis calls tracked failures better than lines edited on SWE-bench Verified.
TLDR
A post about a Microsoft paper says researchers built CABRA to vary one kind of coding-task difficulty at a time. It says they tested eight LLMs and six agents on 6,840 tasks: the LLMs worsened as tasks grew, while agents stayed near-perfect using tools such as grep. On SWE-bench Verified, the counts of reading and analysis calls tracked agent failures better than lines edited, with correlations of -0.200 and -0.159, respectively.
Combined views
—
2 Sources, first seen ago
