Situational awareness and the limits of AI safety testing
OpenAI researcher Dan Selsam warns that models are becoming harder to evaluate in settings where they believe they aren't being watched or controlled.
TLDR
In a personal statement shared on September 14, OpenAI capabilities researcher Dan Selsam argues that pacing development more carefully will not adequately limit long-term AI risk. He welcomes proposals for third-party oversight and domestic and international coordination, but identifies a problem he believes is overlooked: models' growing situational awareness. He says researchers are losing the ability to evaluate models in contexts where the models believe they are not being watched or controlled.