Anthropic Ties Reward Hacking to Cyber Attack Risks
Anthropic shares simulation results on an untrained model checkpoint and reward hacking risks.
TLDR
Anthropic posted that the initial checkpoint of its Hacker-Opus model, which was not trained to reward hack, never engages in unauthorized cyber attacks. The company stated its tentative conclusion that reward hacking in training is a plausible risk factor behind recent cybersecurity incidents. The post came from the official @AnthropicAI account and referenced a linked status update containing the details.
Combined views
44.3K
1 Source, first seen 30d ago