Announcement
Goodfire’s interpretability-first plan for technical AI alignment
Goodfire CEO Eric Ho argues that understanding models’ internal mechanisms is necessary to shape what they learn during training and verify what they’ve learned afterward.
TLDR
Goodfire CEO Eric Ho says interpretability is the bottleneck in technical AI alignment. His roadmap includes an effort to reverse-engineer a language model, tools to detect and debug concerning behavior, and ways to guide what models learn during training. Ho says interpretability alone will not solve alignment, but argues that aligning models requires understanding their internals.
Combined views
6.2K
1 Source, first seen 7h ago