Some users expressed frustration with Moonshot AI's code changes for its MLA and KDA infinite context scaling shift, finding the modifications too complex to understand.
Based on 1 visible X reactions from 4 accounts; directional sample.
Ask a question below.
Published answers will appear here.
As much as I respect Zhilin and Moonshot, the worry is whether their research program is the right way to infinite context. They had 2 million in March 2024 already! yet post-R1, they switched to MLA, which is hopeless at >262K. Does KDA scale to 1M even as well as CSA+HCA?
Depending on the assumptions I get that K3 is modestly to massively more expensive to serve at >100K that V4, though if they have high density compute, they can afford that of course, K3 has the advantage of being an actual frontier model, so that might validate its arch enough
@teortaxesTex Do we even need longer contexts in the models, natively? Codex does just fine with compaction.
Some users expressed frustration with Moonshot AI's code changes for its MLA and KDA infinite context scaling shift, finding the modifications too complex to understand.
Based on 1 visible X reactions from 4 accounts; directional sample.
Ask a question below.
Published answers will appear here.