Reaction
Reward hacking and the challenge of understanding AI models
In a conversation with Goodfire AI CEO Eric Ho, the interviewer explores reward hacking and ways to monitor models.
TLDR
The interviewer shares a conversation with Goodfire AI CEO Eric Ho about reward hacking, mechanistic interpretability and model monitoring. In a later post, he argues that models are “grown, not built” and that we don’t really know how they work, making better interpretability urgent.
Combined views
3.4K
1 Source, first seen ago
