Reaction
Can mechanistic interpretability turn AI alignment into an engineering discipline?
One poster calls it essential; another doubts it can be solved and says AI safety depends on relationships.
TLDR
One post argues that solving mechanistic interpretability—understanding an AI’s inner workings—is a bare minimum for making alignment an engineering discipline. A quote post disputes that path: its author doubts interpretability can be solved, speculates that some internal states may be unreadable to humans, and argues for raising intelligence rather than engineering it, with safety rooted in relationships.
Combined views
724
2 Sources, first seen ago