Letting AI choose what to read cut Gemma-4-31B attention work by 52.0%, a user reports
A post describing Declarative Attention says the model selects the context it needs and the inference engine skips the rest, rather than rereading everything for each new token. The reported savings came with lower accuracy.
TLDR
A user sharing the paper “Language Models Can Control Their Own Attention” describes Declarative Attention as letting models specify which context they need, without a separate scorer searching the whole context first. Across 15 long-context tasks, the user reports attention-work reductions of 52.0% for Gemma-4-31B and 31.1% for Qwen-3.6-27B, alongside accuracy drops of 1.27 and 2.75 percentage points, respectively.
Combined views
1.1K
1 Source, first seen 26d ago