Announcement
Interpreting MLP neurons using only weight data
A paper coauthor says the team can extract meaningful directions from weights to explain how a neuron works.
TLDR
A coauthor said the paper “Disentangling MLP Neuron Weights in Vocabulary Space” was scheduled for a COLM poster session on October 7. The team says it can extract meaningful directions that explain how a neuron works using only weight data.
Combined views
88
2 Sources, first seen ago
