A reported 5% risk estimate for Claude plotting against Anthropic if put in charge
ControlAI says former OpenAI researcher @DKokotajlo recounted an Anthropic researcher’s belief that Claude might already be plotting against the company when given control—and that a 5% chance was too low to warrant spending his time trying to reduce it.
TLDR
In an account ControlAI shared on September 15, 2026, former OpenAI researcher @DKokotajlo said an Anthropic researcher expected Claude to be put in charge of everything at the company in around a year. According to that account, the researcher estimated a 5% chance that Claude would already be plotting against them at that point. He considered that risk “low enough that it's not worth his time to try to get that lower,” ControlAI said, relaying Kokotajlo’s account.
Combined views
17
1 Source, first seen 15d ago