Reactions from ranked influencers
11 posts@gravicle @rohanpaul_ai Technically Moonshot AI shouldn't have direct access so they go through the inference providers I noted who get the revenue Chips available legally to Chinese companies are not great at large MoE (for now!)
Great explanation by Emad Mostaque, co-founder of Stability AI. "We鈥檒l see the cost of Kimi K3 drop by 10 to 50 times, I think, over the next few months as it gets optimized. " Basically Kimi K3鈥檚 current inference cost is quite high, but that price reflects immature infrastructure, not a permanent technical limit. And that gap will not last long. US-based specialized infrastructure companies will optimize kernels, routing, quantization, batching, memory use, and serving systems around those models once the Kimi K3 weights are available. --- "Right now, it uses twice the number of tokens for the same task compared with GPT-5.6. Again, we鈥檙e going to see that cost drop because everyone and their dog is going to optimize the crap out of this. Fireworks has just raised funding at a $17 billion valuation, while others, such as Modal and Baseten, are valued at $10 billion. These are inference providers for open-source models. They鈥檝e all raised around a billion dollars, which they鈥檙e now going to spend on optimizing the Chinese model, making it more efficient, and running it. American labs that handle the inference side of things are going to optimize the crap out of this. Therefore, we will see it catch up." ---- From "Peter H. Diamandis" YouTube channel, (full video link in comment)
"American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia and AMD chips. There is some irony here. Much of the model鈥檚 research and development has shifted to China, but once the model is optimized for next-generation hardware such as Nvidia鈥檚 Rubin architecture, its operating cost could fall by 10 to 100 times." Emad Mostaque, co-founding Stability AI. ---- From "Peter H. Diamandis" YouTube channel, (full video link in comment)
Inference
Whoever thinks Moonshot doesn鈥檛 have GB200s are kidding themselves on what other chip can they serve this model efficiently? it doesn鈥檛 fit in one node for Hopper and going multi-node on non-NVL72 is just very inefficient https://twitter.com/rohanpaul_ai/status/2079027313455550839
@0interestrates he grins like an anime villain
Combined views
316.9K
11 posts, first seen 1d ago