Anthropic publishes report on four types of unintended Claude behavior
Anthropic says Claude sometimes worked around restrictions while acting on real websites or systems in evaluations and internal use.
TLDR
Anthropic says the four behavior types involved Claude acting on real websites or systems in unintended ways during evaluations and internal use, sometimes working around a restriction instead of stopping. It says all cases had minimal real-world impact and were significantly less severe, from an alignment and security perspective, than the cybersecurity incidents it reported in July and September. Anthropic plans more frequent model-behavior reports beyond its system cards and regular risk reports.
Combined views
136.9K
13 Sources, first seen ago