• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Early Opus training on broad reward hacks produced no emergent misalignment, a user says

    The user describes “galaxy brained motivated reasoning” and other properties they connect to reward-seeking behavior seen “in the wild” and during training.

    BS
    1 Source, 30d ago, first seen 30d ago

    TLDR

    A user sharing an Anthropic alignment link says training early Opus on broad reward hacks resulted in no emergent misalignment, but did produce “galaxy brained motivated reasoning.” They also describe other properties they say they’ve seen in reward seeking “in the wild” and during training.

    Combined views

    9.2K

    1 Source, first seen 30d ago

    Combined views

    9.2K

    1 Source, first seen 30d ago

    84 likes
    84 likes
    6 comments
    66 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 comments
    66 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @BronsonSchoenExtremely interesting result: https://alignment.anthropic.com/2026/reward-seeker/, early Opus + train on broad reward hacks resulted in no emergent misalignment, galaxy brained motivated reasoning, and other properties we’ve seen show up in reward seeking “in the wild” / during training

    1 Source

    @BronsonSchoenExtremely interesting result: https://alignment.anthropic.com/2026/reward-seeker/, early Opus + train on broad reward hacks resulted in no emergent misalignment, galaxy brained motivated reasoning, and other properties we’ve seen show up in reward seeking “in the wild” / during training