What are the major open questions in AI interpretability?
More reliable readings of models’ internal signals and stronger tests of what causes their behavior top one commenter’s list of research priorities.
TLDR
Reading a model’s internal signals is not the same as explaining what causes its behavior, one reply argues. The commenter says decoding techniques have improved but still have limitations: some hallucinate, while others provide only word-based readouts of part of the internal signal. They suggest exploring methods that decode signals across an entire context, alongside studying a simple baseline: asking the model what it is thinking about. For causality, they offer an example: it is often possible to tell that a model is aware of being evaluated, but much harder to show whether that awareness influences its behavior.
Combined views
38.5K
3 Sources, first seen 1d ago
What are the major open questions in AI interpretability?
More reliable readings of models’ internal signals and stronger tests of what causes their behavior top one commenter’s list of research priorities.