Real-SWE’s top-scoring AI coding agent reportedly failed over 60% of the time
The New Stack reports on Real-SWE, a benchmark that tests AI coding agents on private, proprietary codebases.
TLDR
The New Stack reports that even the top-scoring AI coding agent failed over 60% of the time on Real-SWE. The benchmark tests models on proprietary codebases.
Real-SWE’s top-scoring AI coding agent reportedly failed over 60% of the time
The New Stack reports on Real-SWE, a benchmark that tests AI coding agents on private, proprietary codebases.
TLDR
The New Stack reports that even the top-scoring AI coding agent failed over 60% of the time on Real-SWE. The benchmark tests models on proprietary codebases.