Report
A proposed learning method uses a growing gradient-signal cache instead of per-step linear-layer updates
A researcher says the paper derives the method from a duality between SGD and linear attention.
TLDR
A researcher says a new paper with a Google PI proposes a learning method derived from the duality between SGD and linear attention. Instead of updating linear layers with gradient descent at every training step, the method maintains a growing “KV cache” of gradient signals to attend over. The researcher says the neural network grows during training.
Combined views
2.4K
2 Sources, first seen ago
65 likes42 saves19 reposts