• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Collinear AI Releases CWE-Bench for Coding Agents

    Nazneen Rajani announces benchmark of tasks testing coding agents on real code defense.

    NR
    1 Source, 28d ago, first seen 28d ago

    TLDR

    Nazneen Rajani, founder and CEO of Collinear AI, released CWE-bench. The benchmark contains 100 held-out audit-and-patch tasks that test whether coding agents can defend real code. Rajani stated the leading agent passes 47 percent of them and 18 tasks remain unsolved by any agent. She added that frontier models on the benchmark show de-correlated errors. The source describes CWE-bench as a frontier benchmark for AI cybersecurity that measures coding agent capabilities on every class of known vulnerability.

    Combined views

    66.7K

    1 Source, first seen 28d ago

    Combined views

    66.7K

    1 Source, first seen 28d ago

    67 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    67 likes
    20 comments
    38 saves
    9 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    20 comments
    38 saves
    9 reposts

    1 Source

    @nazneenrajaniToday we’re releasing CWE-bench: 100 held-out audit-and-patch tasks testing whether coding agents can defend real code. The leading agent passes 47%. 18 tasks remain unsolved by any agent. This means frontier models on our benchmark have de-correlated errors. Blog: https://blog.collinear.ai/p/cwe-bench

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @nazneenrajaniToday we’re releasing CWE-bench: 100 held-out audit-and-patch tasks testing whether coding agents can defend real code. The leading agent passes 47%. 18 tasks remain unsolved by any agent. This means frontier models on our benchmark have de-correlated errors. Blog: https://blog.collinear.ai/p/cwe-bench