Token superposition training reportedly reaches comparable or better loss with about 2.5× less compute in a 10B MoE test · Digg