• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Hugging Face Retweets Sliding Window Attention Resource

    Retweet points readers to Papers with Code page on the attention method.

    HF
    JS
    KA
    6 Sources, 30d ago, first seen 30d ago

    TLDR

    Hugging Face retweeted a post from Niels Rogge that directs people to a Papers with Code entry on Sliding Window Attention. The post responds to questions about the technique. Papers with Code is presented in the post as a site covering trending AI research papers, code, datasets, methods, and evaluation leaderboards. The retweet comes from the official Hugging Face account, which manages the Hub and Transformers library.

    Combined views

    14.3K

    6 Sources, first seen 30d ago

    Combined views

    14.3K

    6 Sources, first seen 30d ago

    144 likes
    144 likes
    10 comments
    58 saves
    63 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    10 comments
    58 saves
    63 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 Sources

    @huggingfaceRT @NielsRogge: For folks wondering what Sliding Window Attention is, there's a method for it on Papers with Code Sliding Window Attention…
    @yacineMTBRT @giffmana: @jm_alexia @JFPuget @RheaSukthanker @CameronPashmina @Emy_Aze I'm a bit disappointed you made the title so misleading. Just a…
    @SimonGoodman_@jm_alexia @RheaSukthanker @CameronPashmina @Emy_Aze ACTUALLY @SchmidhuberAI did it before in https://proceedings.mlr.press/v139/schlag21a you could have at least cited it
    @rohanpaul_aiSo happy to see this new Microsoft paper. If lower inference memory is the goal, this paper finds training-free Sliding Window Attention beats most retrofitted linear-attention methods, making it the simpler default to try first. Keep only a small recent window, plus the first 4 “sink” tokens that models rely on. With a 64-token window, this training-free setup had the best average downstream score in 9 of 11 model comparisons and recovered 99.0% of the full-attention baseline average. Many linear-attention alternatives need additional post-training; this version of SWA needs none. The gap grew on long-context reasoning. At 4K context, SWA reached 17.2%–23.0% on the Needle-in-a-Haystack tasks, while LoLCATs reached at most 5.8%; on BABILong, SWA scored 15% versus 3%. In their speed and memory test, the 64-token SWA setup was fastest and used the least memory. Full attention still wins badly on long context, but for fixed, low memory without retraining, the paper recommends trying SWA with attention sinks first. – arxiv. org/abs/2608.28444 Title: "Sliding-window beats linear attention"
    @SchmidhuberAIRT @SimonGoodman_: @jm_alexia @RheaSukthanker @CameronPashmina @Emy_Aze ACTUALLY @SchmidhuberAI did it before in https://proceedings.mlr.press/v139/schlag21a yo…

    6 Sources

    @huggingfaceRT @NielsRogge: For folks wondering what Sliding Window Attention is, there's a method for it on Papers with Code Sliding Window Attention…
    @yacineMTBRT @giffmana: @jm_alexia @JFPuget @RheaSukthanker @CameronPashmina @Emy_Aze I'm a bit disappointed you made the title so misleading. Just a…
    @SimonGoodman_@jm_alexia @RheaSukthanker @CameronPashmina @Emy_Aze ACTUALLY @SchmidhuberAI did it before in https://proceedings.mlr.press/v139/schlag21a you could have at least cited it
    @rohanpaul_aiSo happy to see this new Microsoft paper. If lower inference memory is the goal, this paper finds training-free Sliding Window Attention beats most retrofitted linear-attention methods, making it the simpler default to try first. Keep only a small recent window, plus the first 4 “sink” tokens that models rely on. With a 64-token window, this training-free setup had the best average downstream score in 9 of 11 model comparisons and recovered 99.0% of the full-attention baseline average. Many linear-attention alternatives need additional post-training; this version of SWA needs none. The gap grew on long-context reasoning. At 4K context, SWA reached 17.2%–23.0% on the Needle-in-a-Haystack tasks, while LoLCATs reached at most 5.8%; on BABILong, SWA scored 15% versus 3%. In their speed and memory test, the 64-token SWA setup was fastest and used the least memory. Full attention still wins badly on long context, but for fixed, low memory without retraining, the paper recommends trying SWA with attention sinks first. – arxiv. org/abs/2608.28444 Title: "Sliding-window beats linear attention"
    @SchmidhuberAIRT @SimonGoodman_: @jm_alexia @RheaSukthanker @CameronPashmina @Emy_Aze ACTUALLY @SchmidhuberAI did it before in https://proceedings.mlr.press/v139/schlag21a yo…