A 5% risk estimate for Claude plotting against Anthropic when put in charge
ControlAI says former OpenAI researcher @DKokotajlo recounted an Anthropic researcher's belief that Claude had a 5% chance of already plotting against the company when put in charge of everything there.
TLDR
On September 15, 2026, ControlAI shared an account from former OpenAI researcher @DKokotajlo's speech at its London event. According to ControlAI, he said an Anthropic researcher expected Claude to be put in charge of everything at the company in around a year, and estimated a 5% chance it would already be plotting against them at that point. In the relayed account, the researcher considered that chance “low enough that it's not worth his time to try to get that lower”.
Combined views
10.9K
1 Source, first seen 15d ago