Compute Shortage Threatens Scaling Of Frontier AI Models Like Kimi K3
Reactions from ranked influencers
2 postsThere are a lot of assumptions baked into whether this continues, and things that can happen to change those assumptions: - Frontier intelligence performance has to continue to scale. If we “tap out” on how good the models get, many will catch up to the frontier and whoever has compute will drive down costs. (b) from above. - If demand for frontier intelligence ceases to grow (“most tasks just don’t need it”), than much cheaper near-frontier models will capture more dollars. (a) from above. - If you can suddenly create and serve frontier intelligence at a radically cheaper cost or compute structure, no one will pay a premium. (c) from above. - If compute was infinitely available, one might argue that the demand for frontier intelligence would decline if significantly if near-frontier intelligence was much cheaper. Related to (a) from above. - Just because more dollars accrue to the frontier does not necessarily imply it is economical. It may well be the case that cost to train, capture talent, compute, generate demand (s&m) etc is so high that even a large amount of dollars in is loss-making. What’s true is that there are many many companies taking bets on many of these assumptions and if they’re right, they will be tremendously valuable companies.
The story of AI in the next few years is going to be compute: an essay on the future of AI. K3 in 2 days is already #10 on OpenRouter with ~140B tok/day, and it’s infra is crumbling. Throughput is down from 30tok/s to 13tok/s, E2E latency is up to 72s and time to first token is >20s! It would cost a minimum of $500k to buy the 8 B300s it would take to serve even quantized Kimi K3 and ~$4M for the more recommended GB300 NVL72 rack. I don’t think Moonshot has the compute available to scale to their demand! In fact, even the US based inference providers will likely not be able to scale capacity as much as they’d like even if they were to host it: a 2.8T model is no joke. GPU providers (neoclouds etc) are doing 3yr and I recently hear 5yr commits with an ungodly 30% down, and customers are chomping it up. Prices continue to go to the moon. The two big labs, hyperscaler clouds, Grok and Meta have compute deals locked in prior, and the rest are fighting for scraps. Tier 1 neoclouds (coreweave/nebius etc) are rumored to not even small “smaller” customers. Meta is the biggest wildcard here. With ~7GW of compute by eoy 2026 and no clear big model ties, they either get to frontier on their own or can host the most Kimi K3 capacity (unless they sell it to the labs). Even though the price of models has fallen over time, it’s worth noting that the price of frontier has not. 3yrs ago, GPT-4 released at $60/M, o1 at $60/M, Opus 4 at $75/M, GPT5 at $10/M, Fable at $50/M and now Sol at $30/M and K3 at $15/M. Even if you consider K3 frontier, that’s only a 4-5x flux in 3yrs. In that time, frontier demand has increased at least 3+ ooms and frontier intelligence performance has gone 32x at least by task time by METR. Essentially, so long as a) the demand for frontier intelligence continues to grow to near infinity, b) the frontier continues to grow in performance, even as c) if the price of frontier declines a little, the value accrued to frontier grows significantly! And there’s a tremendous bull case for those who have locked up compute if you’re bitter lesson pilled and believe larger models will always be smarter models.
Combined views
166.7K
2 posts, first seen 21h ago