• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Tanishq Abraham Shares SMELT on Looped MoE Transformers

    Tanishq Mathew Abraham posts ablations on looping Mixture-of-Experts transformers with matched compute.

    JL
    T(
    TM
    7 Sources, 29d ago, first seen 29d ago

    TLDR

    Tanishq Mathew Abraham posted about the SMELT paper. The work examines looping on Mixture-of-Experts Transformers while matching per-token FLOPs, total non-embedding parameters, and KV cache. A series of ablations leads to the SMELT recipe, described as Sparse MoE Transformer where middle layers loop twice. The post includes the paper title and a portion of its abstract. The linked source on arXiv.org presents the same title and expands on the evaluation approach for looped transformers.

    Combined views

    130.2K

    7 Sources, first seen 29d ago

    Combined views

    130.2K

    7 Sources, first seen 29d ago

    989 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    989 likes
    17 comments
    651 saves
    108 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    17 comments
    651 saves
    108 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 Sources

    @iScienceLuvrSMELT: Scaling Laws for Compute-Matched MoE Looped Transformers "We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT’s loss drops faster with compute, saving 6.8–18.0% of training FLOPs on the compute-optimal frontier." paper link: https://arxiv.org/abs/2609.01343
    @teortaxesTex> middle layers loop twice The same factor as in Loop the Loopies! (yes really) from IQuest Research, they scaled to 20B A2B. But Loopies repeats individual layers; SMELT repeats on block level. No ablation of layer-wise looping, sadly.
    @wangsw5653Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: https://arxiv.org/abs/2609.01343
    @scaling01China bros will pounce and again this shows how absolutely underrated ByteDance Seed is as a lab
    @jasondeanleeRT @wangsw5653: Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws…

    7 Sources

    @iScienceLuvrSMELT: Scaling Laws for Compute-Matched MoE Looped Transformers "We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT’s loss drops faster with compute, saving 6.8–18.0% of training FLOPs on the compute-optimal frontier." paper link: https://arxiv.org/abs/2609.01343
    @teortaxesTex> middle layers loop twice The same factor as in Loop the Loopies! (yes really) from IQuest Research, they scaled to 20B A2B. But Loopies repeats individual layers; SMELT repeats on block level. No ablation of layer-wise looping, sadly.
    @wangsw5653Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: https://arxiv.org/abs/2609.01343
    @scaling01China bros will pounce and again this shows how absolutely underrated ByteDance Seed is as a lab
    @jasondeanleeRT @wangsw5653: Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws…