Real-SWE’s top-scoring coding agent reportedly failed over 60% of the time
The New Stack describes Real-SWE, a benchmark that tests AI coding models on proprietary codebases.
TLDR
AI coding agents struggled on Real-SWE’s proprietary-code tests, The New Stack reports, with even the benchmark’s top scorer failing over 60% of the time.
Combined views
440
1 Source, first seen 8h ago
Real-SWE’s top-scoring coding agent reportedly failed over 60% of the time
The New Stack describes Real-SWE, a benchmark that tests AI coding models on proprietary codebases.
TLDR
AI coding agents struggled on Real-SWE’s proprietary-code tests, The New Stack reports, with even the benchmark’s top scorer failing over 60% of the time.