Claude-generated methods reportedly beat human-proposed baselines on seven alignment failures
Surge AI, which ran the human-researcher baseline for Anthropic’s study, highlights a key caveat: the agents could iterate repeatedly, while each human researcher submitted one idea.
TLDR
Surge AI describes Anthropic research using Claude agents to search the literature, propose alignment methods, train models, evaluate results and iterate. According to Surge, the agents found methods that improved safety benchmarks while preserving general capabilities across 10 alignment failures, including deception, jailbreaks and privacy violations. The strongest methods also generalized to models up to 4.7 times larger than those they optimized against.
Surge says it ran the human baseline with 28 experienced technical AI safety researchers, each given up to eight hours to propose a method. The automated researchers’ methods outperformed human-proposed baselines on the seven failures where human ideas were collected. Surge notes Anthropic’s caution about that comparison: agents could iterate repeatedly, while humans submitted one idea each.
Combined views
10
1 Source, first seen 15d ago