Reaction
Linear substructures as a possible lens on how large AI models 'learn how to learn'
A user says Sheridan's attention-bundle OV technique is closely related to the Jacobian lens, describing a linear relation decomposed along one causal axis.
TLDR
A user suggests linear substructures seem to emerge from massive training and might help explain some of the ways large models 'learn how to learn.' They also connect Sheridan's attention-bundle OV technique to the Jacobian lens and link to a paper they say offers a different decomposition.
Combined views
103
1 Source, first seen 7h ago