Matryoshka Attribution debuts with a claimed No. 1 result on the Mechanistic Interpretability Benchmark
The paper’s authors say MAttr uses gradient descent to find which parts of a neural network are responsible for a behavior. They report a benchmark result of 2.9× the runner-up.
TLDR
The authors introduce Matryoshka Attribution (MAttr), a method that uses gradient descent—an optimization technique—to identify neural-network components responsible for a behavior. They report that it ranks first on the Mechanistic Interpretability Benchmark, at 2.9× the runner-up. A coauthor says MAttr learns a nested ranking of model components across sparsity levels, directly optimizes intervention performance, transfers across tasks and extends to parameter attribution.
Combined views
6.1K
3 Sources, first seen 9h ago
Matryoshka Attribution debuts with a claimed No. 1 result on the Mechanistic Interpretability Benchmark
The paper’s authors say MAttr uses gradient descent to find which parts of a neural network are responsible for a behavior. They report a benchmark result of 2.9× the runner-up.