Paper Proposes Declarative Attention for Language Models
Tweet highlights arXiv paper on language models managing attention via a new protocol.
TLDR
Elvis Saravia posted about an arXiv paper titled Language Models Can Control Their Own Attention. The work by Namgyu Ho at KAIST AI along with Tal Schuster and Cicero Nogueira dos Santos at Google DeepMind describes how models reread their full KV cache for each generated token. It introduces Declarative Attention as a protocol for models to handle their own attention. The post links the arXiv entry and a dair.ai summary of the paper.
Combined views
65.5K
6 Sources, first seen 27d ago