Real-SWE's top-scoring AI coding agent reportedly failed over 60% of the time
The New Stack reports that Real-SWE tests AI models on private, proprietary codebases.
TLDR
The New Stack reports that even the top scorer on Real-SWE failed over 60% of the time. The benchmark tests AI coding models on proprietary codebases.
Combined views
1.2K
2 Sources, first seen ago
likes