2 stories tagged by Digg
AI
Goodfire says sparse autoencoders are useful for exploring a model when you don’t know what to look for. For a specific target, such as reward hacking, it recommends a probe.
The AI interpretability company announces access to its research tool.