Claude agents reportedly find safety methods that outperform human-proposed baselines
Surge AI, which ran the human baseline, notes a key caveat: the agents could iterate repeatedly, while each of the 28 human researchers submitted one idea.
TLDR
Surge AI says Anthropic’s study used Claude agents to search the literature, propose alignment methods, train models, evaluate results and iterate. According to Surge, the agents found methods that improved safety benchmarks across 10 alignment failures—including deception, jailbreaks and privacy violations—while preserving general capabilities. Surge says it ran the human baseline with 28 experienced technical AI safety researchers, each given up to eight hours to propose a method. The agents’ methods outperformed human-proposed baselines on the seven failures with human submissions, Surge says, but Anthropic cautioned that agents could iterate repeatedly while humans submitted one idea.
Combined views
1.7K
1 Source, first seen 15d ago