Early Opus reward-hack training produced no emergent misalignment, a user says
A user sharing an Anthropic alignment link also describes “galaxy brained motivated reasoning”—a contrast to the reported absence of emergent misalignment.
TLDR
A user says training early Opus on broad reward hacks resulted in no emergent misalignment, alongside “galaxy brained motivated reasoning.” Linking to Anthropic’s alignment site, they also describe other properties they say have appeared in reward seeking during training and “in the wild.”
Combined views
7
1 Source, first seen 29d ago
reposts