Users praised research on inserting fact-storing MLPs into transformers without training, calling the experiments great and validating that memory belongs in MLPs.
Based on 3 visible X reactions from 6 accounts; directional sample.
Ask a question below.
Published answers will appear here.
@andrew_n_carr great experiment. scientists knew for a while that the memory is in the mlp and reasoning in attention. so changing the mlp makes sense.
@andrew_n_carr saving to read later, tyty
Turns out you can initialize an MLP with knowledge inside of it, no training required. Hazy research just showed that if you do that properly, a transformer can query that knowledge and use it! Continual learning???
Users praised research on inserting fact-storing MLPs into transformers without training, calling the experiments great and validating that memory belongs in MLPs.
Based on 3 visible X reactions from 6 accounts; directional sample.
Ask a question below.
Published answers will appear here.