New Paper Unifies Interpretability Methods via Tensor Product Representations
Enyan Zhang derives four mechanistic interpretability methods from tensor product representations.
TLDR
Enyan Zhang announced a new paper titled A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability. The work states it derives four popular interpretability methods from the single hypothesis of Tensor Product Representations. Tom McCoy, assistant professor at Yale, quoted the thread and called the paper exciting for mechanistic interpretability researchers. He noted the unification under the TPR formalism and tagged collaborators including Paul Smolensky and Tal Linzen. The post includes an overview slide attachment.
Combined views
18.5K
4 Sources, first seen 26d ago