• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Shared memory is claimed to improve looped Transformers while cutting memory use

    The researchers say later recursions read the first recursion’s key-value cache while keeping a short window of their own.

    CLSCL
    Giovanni MoneaGM
    2 Sources, ,

    TLDR

    In a preprint, the researchers report that training looped Transformers to share memory improved quality rather than hurting it. Their hybrid model with five recursions had 1.12–1.82 lower validation perplexity on FineWeb-Edu than a same-size standard Transformer while using 76–79% less context memory, they say.

    Combined views

    3K

    2 Sources, first seen 2h ago

    Combined views

    3K

    2 Sources, first seen 2h ago

    55 likes
    2h ago
    first seen 2h ago
    55 likes
    2 comments
    34 saves
    13 reposts
    2 comments
    34 saves
    13 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Giovanni Monea@giomonea⚠️ Stop pretraining your looped Transformer with a separate KV cache per recursion! In our recent preprint, we show that *memory sharing* not only saves memory but is also a net-positive inductive bias! Less memory, same flops, higher quality. 🔗 https://arxiv.org/abs/2610.02383 🧵2h
    CLS@ChengleiSiRT @giomonea: ⚠️ Stop pretraining your looped Transformer with a separate KV cache per recursion! In our recent preprint, we show that *me…2h

    2 Sources

    Giovanni Monea@giomonea⚠️ Stop pretraining your looped Transformer with a separate KV cache per recursion! In our recent preprint, we show that *memory sharing* not only saves memory but is also a net-positive inductive bias! Less memory, same flops, higher quality. 🔗 https://arxiv.org/abs/2610.02383 🧵2h
    CLS@ChengleiSiRT @giomonea: ⚠️ Stop pretraining your looped Transformer with a separate KV cache per recursion! In our recent preprint, we show that *me…2h