Reaction
Mechanistic interpretability as a proposed prerequisite for engineering AI alignment
One commenter argues that solving mechanistic interpretability is the bare minimum for making AI alignment an engineering discipline.
TLDR
One commenter argues that AI alignment needs mechanistic interpretability solved before it can become an engineering discipline. They contrast that goal with what they call “superintelligent animal husbandry.”
Combined views
23.4K
4 Sources, first seen ago