• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Recirculation Technique Recycles Activations in Transformers

    The method reuses top-layer activations in lower layers during inference on existing models.

    RA
    TK
    RL
    18 Sources, 43d ago, first seen 43d ago

    TLDR

    Researcher Michael C. Mozer posted a paper describing recirculation, an enhancement for off-the-shelf foundation models. The approach passes activations from higher layers into lower layers during subsequent inference steps. It requires no retraining. The paper states the change markedly reduces perplexity and boosts accuracy on generation and reasoning tasks while adding near-zero latency. Several researchers shared the work on X and highlighted its simplicity.

    Combined views

    431.2K

    18 Sources, first seen 43d ago

    Combined views

    431.2K

    18 Sources, first seen 43d ago

    2.4K likes
    2.4K likes
    65 comments
    2.3K saves
    280 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    65 comments
    2.3K saves
    280 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    18 Sources

    @rosinalityhttps://arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (https://arxiv.org/abs/2608.08888). Why does this work without training?
    @yacinelearninglike everything works in ai now you can literally do whatever
    @SonglinYang4RT @rosinality: https://arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the nex…
    @DanielleFongRT @yacinelearning: like everything works in ai now you can literally do whatever
    @mc_mozerWhat if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—without retraining? What if that tweak incurred near zero latency cost during generation and supported indefinite state tracking? https://arxiv.org/abs/2608.17981
    @TheGradientRT @mc_mozer: What if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—with…
    @techdreamergc@mc_mozer Interesting distinction from chain-of-thought. CoT is the default for reasoning gains, but state tracking feels more fundamental. If feedforward depth limits state updates, inference-time recurrence is a natural fix.
    @scaling01like it wouldn't surprise me if OpenAI or Ant or both have something like this:
    @_arohan_The OG posts a banger
    @anirudhg9119Pre-trained Transformers may have latent architectural affordances that training never explicitly asked for. Cool work.

    18 Sources

    @rosinalityhttps://arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (https://arxiv.org/abs/2608.08888). Why does this work without training?
    @yacinelearninglike everything works in ai now you can literally do whatever
    @SonglinYang4RT @rosinality: https://arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the nex…
    @DanielleFongRT @yacinelearning: like everything works in ai now you can literally do whatever
    @mc_mozerWhat if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—without retraining? What if that tweak incurred near zero latency cost during generation and supported indefinite state tracking? https://arxiv.org/abs/2608.17981
    @TheGradientRT @mc_mozer: What if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—with…
    @techdreamergc@mc_mozer Interesting distinction from chain-of-thought. CoT is the default for reasoning gains, but state tracking feels more fundamental. If feedforward depth limits state updates, inference-time recurrence is a natural fix.
    @scaling01like it wouldn't surprise me if OpenAI or Ant or both have something like this:
    @_arohan_The OG posts a banger
    @anirudhg9119Pre-trained Transformers may have latent architectural affordances that training never explicitly asked for. Cool work.