Tanishq Abraham Shares SMELT on Looped MoE Transformers
Tanishq Mathew Abraham posts ablations on looping Mixture-of-Experts transformers with matched compute.
TLDR
Tanishq Mathew Abraham posted about the SMELT paper. The work examines looping on Mixture-of-Experts Transformers while matching per-token FLOPs, total non-embedding parameters, and KV cache. A series of ablations leads to the SMELT recipe, described as Sparse MoE Transformer where middle layers loop twice. The post includes the paper title and a portion of its abstract. The linked source on arXiv.org presents the same title and expands on the evaluation approach for looped transformers.
Combined views
130.2K
7 Sources, first seen 29d ago