Report
Reward hacking and the search for a clearer view inside AI models
An interviewer shares a conversation with Goodfire AI CEO Eric Ho about reward hacking, model monitoring and interpretability.
TLDR
The interviewer shares a conversation with Goodfire AI CEO Eric Ho about reward hacking and ways to examine what happens inside AI models. The post’s chapter list raises questions about whether models know when they’re cheating, how to monitor them and whether neural networks could be decoded by 2028.
Combined views
2.1K
2 Sources, first seen 5h ago
