Choosing between sparse autoencoders and probes to interpret AI models
Goodfire says sparse autoencoders are useful for exploring a model when you don’t know what to look for. For a specific target, such as reward hacking, it recommends a probe.
TLDR
Goodfire’s guide describes sparse autoencoders (SAEs) as tools for examining concepts in a model’s activity. It says they work best for open-ended tasks, such as investigating what changed during training. When the goal is to detect a known behavior, Goodfire recommends probes, saying SAEs lose some information and don’t have clearly labeled features by default. The guide also covers choices involved in training an SAE.
Combined views
20.9K
8 Sources, first seen 13h ago
