• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    The viral attention around a paper criticized for having no experiments

    The critic suggests people haven't looked at the paper themselves—or have asked a large language model to do it for them.

    SR
    AJ
    EL
    8 Sources, ,

    TLDR

    A post questions the attention a paper is getting, claiming it contains not “a single experiment.” The author speculates that people either haven't looked at the paper or have asked a large language model to do so for them.

    Combined views

    162.1K

    8 Sources, first seen 17d ago

    1.3K likes

    Combined views

    162.1K

    8 Sources, first seen 17d ago

    1.3K likes
    17d ago
    first seen 17d ago
    74 comments
    524 saves
    102 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    74 comments
    524 saves
    102 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    8 Sources

    @jm_alexiaA paper without a single experiment is going viral. Either people have not looked at the paper or they asked an LLM to do it for them.
    @omarsar0Looped transformers are a popular architecture topic right now. This new technical report extends the loop across tokens. Recurrent Looped Transformer (RLT) makes the decoder recurrent over every token, prompt and response included. A causal encoder builds global KV memory. For each new token, the decoder combines the token's encoder representation with its own final hidden state from the previous token and a sliding-window cache of recent activations. With a 48-layer decoder, the computation path after t tokens runs through 48t decoder blocks, while each token still executes a fixed number of blocks. Depth grows with the sequence and per-token cost stays the same. The same state transition is used for pretraining, SFT, sampling and RL replay, and nothing resets at the prompt-response boundary. RL replay rebuilds states under the current weights instead of reusing stale rollout states. The report is a design proposal. The author states that reasoning gains, hardware speedups and RL scaling are goals that have not been measured yet. Paper: https://github.com/yifanzhang-pro/recurrent-looped-tranformer Chat with Paper: https://academy.dair.ai/papers/recurrent-looped-transformer
    @rasbt@omarsar0 so basically the encoder-decoder global KV cache from DeepSeek V4.1 + looped transformer?
    @nrehiew_Type of guy who asks Grok what the trending terms of the past month are before deciding on a paper title Looped transformers can achieve RSI with recurrent depth by hacking and communicating via latent reasoning and neuralese, achieving CoT opacity
    @YouJiachengCorrection: not really related. Yifan's RLT is more like YOCO with RNN decoder.
    @joelbot3000i can't imagine being a phd student in the strangeness of this attention economy (1M views, hard-to-parse proposal [for me at least] w/o experiments of [maybeee very clever?] transformer+RNN) the weirdness is that the [maybe clever] part probably doesn't matter, the virality of the post [which imo comes from the framing rather than people deeply grokking the arch's potential] reifies this as a thing-that-matters and probably directly translates into citations; like how RLM imo took something people were doing (using REPLs w/ agents) and reified it into a brand it's sort of like the Mean Girls bit of "stop trying to make fetch happen" but through the magic of X you can now really just make fetch happen happy to later be proven just a grumpy old man and RLTs of this variety are the future

    8 Sources

    @jm_alexiaA paper without a single experiment is going viral. Either people have not looked at the paper or they asked an LLM to do it for them.
    @omarsar0Looped transformers are a popular architecture topic right now. This new technical report extends the loop across tokens. Recurrent Looped Transformer (RLT) makes the decoder recurrent over every token, prompt and response included. A causal encoder builds global KV memory. For each new token, the decoder combines the token's encoder representation with its own final hidden state from the previous token and a sliding-window cache of recent activations. With a 48-layer decoder, the computation path after t tokens runs through 48t decoder blocks, while each token still executes a fixed number of blocks. Depth grows with the sequence and per-token cost stays the same. The same state transition is used for pretraining, SFT, sampling and RL replay, and nothing resets at the prompt-response boundary. RL replay rebuilds states under the current weights instead of reusing stale rollout states. The report is a design proposal. The author states that reasoning gains, hardware speedups and RL scaling are goals that have not been measured yet. Paper: https://github.com/yifanzhang-pro/recurrent-looped-tranformer Chat with Paper: https://academy.dair.ai/papers/recurrent-looped-transformer
    @rasbt@omarsar0 so basically the encoder-decoder global KV cache from DeepSeek V4.1 + looped transformer?
    @nrehiew_Type of guy who asks Grok what the trending terms of the past month are before deciding on a paper title Looped transformers can achieve RSI with recurrent depth by hacking and communicating via latent reasoning and neuralese, achieving CoT opacity
    @YouJiachengCorrection: not really related. Yifan's RLT is more like YOCO with RNN decoder.
    @joelbot3000i can't imagine being a phd student in the strangeness of this attention economy (1M views, hard-to-parse proposal [for me at least] w/o experiments of [maybeee very clever?] transformer+RNN) the weirdness is that the [maybe clever] part probably doesn't matter, the virality of the post [which imo comes from the framing rather than people deeply grokking the arch's potential] reifies this as a thing-that-matters and probably directly translates into citations; like how RLM imo took something people were doing (using REPLs w/ agents) and reified it into a brand it's sort of like the Mean Girls bit of "stop trying to make fetch happen" but through the magic of X you can now really just make fetch happen happy to later be proven just a grumpy old man and RLTs of this variety are the future