• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Claude-generated methods reportedly beat human-proposed baselines on seven alignment failures

    Surge AI, which ran the human-researcher baseline for Anthropic’s study, highlights a key caveat: the agents could iterate repeatedly, while each human researcher submitted one idea.

    EC
    1 Source, 15d ago, first seen 15d ago

    TLDR

    Surge AI describes Anthropic research using Claude agents to search the literature, propose alignment methods, train models, evaluate results and iterate. According to Surge, the agents found methods that improved safety benchmarks while preserving general capabilities across 10 alignment failures, including deception, jailbreaks and privacy violations. The strongest methods also generalized to models up to 4.7 times larger than those they optimized against.

    Surge says it ran the human baseline with 28 experienced technical AI safety researchers, each given up to eight hours to propose a method. The automated researchers’ methods outperformed human-proposed baselines on the seven failures where human ideas were collected. Surge notes Anthropic’s caution about that comparison: agents could iterate repeatedly, while humans submitted one idea each.

    Combined views

    10

    1 Source, first seen 15d ago

    Combined views

    10

    1 Source, first seen 15d ago

    1 reposts
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @echenRT @HelloSurgeAI: Anthropic recently published new work on automated alignment researchers: Claude agents that search the literature, propo…

    1 Source

    @echenRT @HelloSurgeAI: Anthropic recently published new work on automated alignment researchers: Claude agents that search the literature, propo…