Collinear AI Releases CWE-Bench for Coding Agents
Nazneen Rajani announces benchmark of tasks testing coding agents on real code defense.
TLDR
Nazneen Rajani, founder and CEO of Collinear AI, released CWE-bench. The benchmark contains 100 held-out audit-and-patch tasks that test whether coding agents can defend real code. Rajani stated the leading agent passes 47 percent of them and 18 tasks remain unsolved by any agent. She added that frontier models on the benchmark show de-correlated errors. The source describes CWE-bench as a frontier benchmark for AI cybersecurity that measures coding agent capabilities on every class of known vulnerability.
Combined views
66.7K
1 Source, first seen 28d ago