@scaling01 @zephyr_z9 V4 GA *is* being delayed, he wanted it in June. Wenfeng hoped to start training a model with 150B active by end of 2026 - April 2027.
@scaling01 @zephyr_z9 It's probably grating to him that K3 or even GLM 5.2 are already useful for self-development. But again: "the first goal of the models we build isn't that users find them good to use, but that we find them good to use ourselves" explains a lot, really
@scaling01 @zephyr_z9 Wenfeng: «China has started to do high quality data annotation only in the last half year» I've been talking of this. Reduction in jaggedness with 5.2 and K3 are emblematic. (Distillation covers the rest)