Attention dilution in long AI context windows
Weaviate Podcast describes a clip on how softmax sums across context tokens, potentially making relevant documents harder to spot.
TLDR
Weaviate Podcast says Siddharth explains that attention’s softmax denominator sums over every context token while a relevant token’s numerator stays the same size, flattening scores as context grows. The post describes BlockSearch as using length-aware scaling and dropping low-scoring documents before attention runs. It describes REFRAG as letting a decoder read precomputed chunk embeddings instead of most raw tokens, then expanding only chunks that need full detail.
Combined views
410
2 Sources, first seen 9h ago