The release of Kimi K3 would be a good time to release GPT-OSS-2, although I don't believe it will happen.
The real Open AI currently comes from China, and that won't change for the time being.
That should at least give us something to think about.…
Secondly, although Kimi Delta Attention has up to 10× lower networking requirements for KV-cache transfers, its large weights require even more network bandwidth to implement an optimization called WideEP, which spreads the weights across different GPUs. 3/8🧵
Furthermore, since the weights occupy more than 1.5 TB of HBM capacity, the KV cache for K3’s KDA and Gated MLA will need to be offloaded to CPU DDR5 and NVMe, even at relatively low user concurrency, because little space remains in HBM. 6/8🧵
Lastly, Jevons’ Paradox means that making attention more efficient will drive wider AI adoption, which will ultimately require more GPUs, HBM, DRAM, and networking—not less. 8/8🧵