Distinguishing AI Safety Gains From Deception Remains Difficult
Matt Yglesias post on evaluation challenges, retweeted by Daniel Kokotajlo.
TLDR
Daniel Kokotajlo, an AI researcher who resigned from OpenAI and now leads the AI Futures Project, retweeted a post by Matt Yglesias. The post states it is extremely hard to tell the difference between we are getting better at preventing misbehavior and the models are g. The packet records the retweet and the quoted text as shared content in the conversation. No further confirmation or independent sources appear in the supplied lines.
Combined views
175
1 Source, first seen 32d ago