• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    Attention dilution in long AI context windows

    Weaviate Podcast describes a clip on how softmax sums across context tokens, potentially making relevant documents harder to spot.

    CS
    WP
    2 Sources, ,

    TLDR

    Weaviate Podcast says Siddharth explains that attention’s softmax denominator sums over every context token while a relevant token’s numerator stays the same size, flattening scores as context grows. The post describes BlockSearch as using length-aware scaling and dropping low-scoring documents before attention runs. It describes REFRAG as letting a decoder read precomputed chunk embeddings instead of most raw tokens, then expanding only chunks that need full detail.

    Combined views

    410

    2 Sources, first seen 9h ago

    Combined views

    410

    2 Sources, first seen 9h ago

    4 likes
    9h ago
    first seen 9h ago
    4 likes
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @weaviatepodcastLong context windows have a math problem that lives in the softmax denominator. 🧮 In this clip, Siddharth explains how attention's softmax denominator sums over every token in context while the relevant token's numerator stays the same size, so attention scores flatten until the right document barely stands out. 🫠 BlockSearch fights this with length-aware softmax scaling and by dropping low-scoring documents before attention runs. 🔥 REFRAG takes a related path: the decoder reads pre-computed chunk embeddings in place of most raw tokens, and a lightweight policy expands only the chunks that need full detail. Both shrink what attention has to look at. 🔎 This chapter discusses attention dilution, the current state of sparse attention, and exciting directions with the help of vector databases!👇 https://www.youtube.com/watch?v=lzw6iFGB2Fg
    @CShorten30RT @weaviatepodcast: Long context windows have a math problem that lives in the softmax denominator. 🧮 In this clip, Siddharth explains ho…
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @weaviatepodcastLong context windows have a math problem that lives in the softmax denominator. 🧮 In this clip, Siddharth explains how attention's softmax denominator sums over every token in context while the relevant token's numerator stays the same size, so attention scores flatten until the right document barely stands out. 🫠 BlockSearch fights this with length-aware softmax scaling and by dropping low-scoring documents before attention runs. 🔥 REFRAG takes a related path: the decoder reads pre-computed chunk embeddings in place of most raw tokens, and a lightweight policy expands only the chunks that need full detail. Both shrink what attention has to look at. 🔎 This chapter discusses attention dilution, the current state of sparse attention, and exciting directions with the help of vector databases!👇 https://www.youtube.com/watch?v=lzw6iFGB2Fg
    @CShorten30RT @weaviatepodcast: Long context windows have a math problem that lives in the softmax denominator. 🧮 In this clip, Siddharth explains ho…