Positive users express excitement about scaling recurrent projectors and iterative dense layers for MoE Transformers, while negative users criticize the Chinese researchers involved for lacking taste and originality.
Based on 2 visible X reactions from 4 accounts; directional sample.
Ask a question below.
Published answers will appear here.
Interesting. Back in early 2024 I was playing with recurrent projectors and iterative application of dense layers (feeding previous output back as input and reusing weights). Saw some potential for efficiency on sequences, but never had the time or compute to scale it properly. Cool to see the idea being pushed at this level with MoE.
@teortaxesTex The only problem with those Chynese quants is they just have no taste. They have absolutely no taste. And I don't mean that in a small way, I mean that in a big way, in the sense that they don't think of original ideas, and they don't bring much culture into their papers
…Pretty wild thing to say but I appreciate the response
Positive users express excitement about scaling recurrent projectors and iterative dense layers for MoE Transformers, while negative users criticize the Chinese researchers involved for lacking taste and originality.
Based on 2 visible X reactions from 4 accounts; directional sample.
Ask a question below.
Published answers will appear here.