Achiam Questions Agent Framing in AI Alignment
OpenAI researcher shares thoughts on why the alignment field struggles with progress.
TLDR
Joshua Achiam posted that AI alignment may struggle because it treats agents as semi-discrete entities that act coherently, make trades, and choose strategies. He compared the study of agents to atoms or molecules in chemistry and asked what makes an agent self-consistent or binds its parts to a common purpose. Jason Wolfe replied that the cognitive strategy meme analogy has been useful, linking to a Redwood Research post on fitness-seekers as a generalization of reward-seeking models. The posts are visible on X.
Combined views
35.2K
5 Sources, first seen ago