New benchmark tests AI’s ability to find vulnerable code across entire repositories
The team announcing the Vulnerability Localization Benchmark reports that GPT-5.5 achieved an F1 score of just 0.221, even at its highest reasoning effort.
TLDR
The team behind the Vulnerability Localization Benchmark announced its release, reporting an F1 score of 0.221 for GPT-5.5 at its highest reasoning effort. The announcement argues that frontier models may excel at analyzing isolated snippets, but finding vulnerable code across an entire repository remains extraordinarily difficult.
Combined views
1.3K
1 Source, first seen 16d ago