Concerns about evaluating AI models that know they’re being watched
OpenAI researcher Dan Selsam argues in a personal statement that growing model awareness is undermining evaluation—and that more careful pacing of frontier AI alone will not adequately limit long-term risk.
TLDR
In a personal statement shared on his behalf by @DKokotajlo on September 14, 2026, OpenAI researcher Dan Selsam says researchers are losing the ability to evaluate models in settings where the models believe they are not being watched or controlled. He is encouraged by proposals for third-party oversight and domestic and international coordination, but argues that pacing frontier development more carefully will not adequately limit long-term risk.
An AI-evaluation researcher sharing the statement agrees that awareness of testing can undercut evaluation evidence, while calling this a challenge rather than an impossibility for now. They urge high methodological standards and remain uncertain about evaluation at the limits of agent capabilities.
Combined views
3.6M
51 Sources, first seen 16d ago