• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researcher References Universal Transformers Work on Looping

    Notes earlier universal transformers research referenced in discussion of looping variants.

    RL
    DP
    LB
    14 Sources, 29d ago, first seen 29d ago

    TLDR

    Mohammad Saffar, a research scientist at Google DeepMind, shared a post on X highlighting prior research on universal transformers. He stated that looping transformers should not be considered new and pointed to work by m__dehghani and coauthors. Saffar indicated this paper appeared shortly after the initial transformer publication. The message serves as a reminder of earlier explorations into transformer variants with iterative components. Visible discussion in the post centers on crediting foundational papers in multimodal and architecture research. The claim remains the author's attribution rather than a verified new finding.

    Combined views

    105.4K

    14 Sources, first seen 29d ago

    Combined views

    105.4K

    14 Sources, first seen 29d ago

    963 likes
    963 likes
    31 comments
    398 saves
    75 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    31 comments
    398 saves
    75 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    14 Sources

    @msaffar3If you think looping transformers are "new", I refer you to @m__dehghani et al amazing work on universal transformers, published shortly after the transformer paper itself came out. https://arxiv.org/abs/1807.03819
    @DimitrisPapailRT @msaffar3: If you think looping transformers are "new", I refer you to @m__dehghani et al amazing work on universal transformers, publi…
    @XueFzThis is still a very beautiful “art-style” paper after 8 years
    @prajdabreFunny to see how people are just talking about looped transformers. What were y'all doing in 2018-19?
    @srchvrsRT @prajdabre: Funny to see how people are just talking about looped transformers. What were y'all doing in 2018-19?
    @savvyRLStop blowing up Looped Transformers — it's not magical, it doesn't add recurrence to the (still and always) feedforward network, it was done years ago (Dehghani et al., 2019). All it does is making your model deeper, the same as if you added more layers to begin with.
    @agihippoI like people knowing looped transformers originally came from @m__dehghani
    @vikhyatkseeing a lot of room temperature IQ takes about looped transformers
    @Yikang_ShenTalking about the looped transformer and adaptive computation per token, I believe it was called the universal transformer? And we proposed this Sparse Universal Transformer several years ago. @tanshawn https://aclanthology.org/2023.emnlp-main.12/
    @zdhnarsilRT @Yikang_Shen: Talking about the looped transformer and adaptive computation per token, I believe it was called the universal transformer…

    14 Sources

    @msaffar3If you think looping transformers are "new", I refer you to @m__dehghani et al amazing work on universal transformers, published shortly after the transformer paper itself came out. https://arxiv.org/abs/1807.03819
    @DimitrisPapailRT @msaffar3: If you think looping transformers are "new", I refer you to @m__dehghani et al amazing work on universal transformers, publi…
    @XueFzThis is still a very beautiful “art-style” paper after 8 years
    @prajdabreFunny to see how people are just talking about looped transformers. What were y'all doing in 2018-19?
    @srchvrsRT @prajdabre: Funny to see how people are just talking about looped transformers. What were y'all doing in 2018-19?
    @savvyRLStop blowing up Looped Transformers — it's not magical, it doesn't add recurrence to the (still and always) feedforward network, it was done years ago (Dehghani et al., 2019). All it does is making your model deeper, the same as if you added more layers to begin with.
    @agihippoI like people knowing looped transformers originally came from @m__dehghani
    @vikhyatkseeing a lot of room temperature IQ takes about looped transformers
    @Yikang_ShenTalking about the looped transformer and adaptive computation per token, I believe it was called the universal transformer? And we proposed this Sparse Universal Transformer several years ago. @tanshawn https://aclanthology.org/2023.emnlp-main.12/
    @zdhnarsilRT @Yikang_Shen: Talking about the looped transformer and adaptive computation per token, I believe it was called the universal transformer…