Reactions from ranked influencers
2 postsKey quotes from the Liang Wenfeng investor presentation: 1. He's convinced that current AI models can't reach AGI and we need continual learning: "Just like CoT—after CoT reached its ceiling, it already surpassed the most top-tier humans, in doing math olympiad problems and writing programs... But it still stops there, that technology couldn’t reach AGI. So you see, AI’s intelligence trajectory is traceable. [A]fter Agent, we think the problem to be solved should be continuous learning—that is, how to let the model learn continuously, rather than you giving it strong training—it should be able to do relatively long-term continuous learning like a human." Observation: This is very similar to how Google views AI progress and next steps (except that Liang does not care for world models and believes that robotics should come at a much later step; more on that below). 2. He thinks that the right order in which to tackle these problems is: continual learning, then RSI (which is a "gradual process", NOT an intelligence explosion), then embodied AI (robotics): "After continuous learning, we might arrive at a singularity. This singularity is: when this model can learn continuously, it can already do all the things humans can do. It could then develop its own version, could research on its own, then develop its own next version, developing more advanced AI models. So it would reach a singularity, able to achieve its own iteration. But this singularity—it’s not really a singularity, it’s also a gradual process. This process might also be a relatively long gradual change, not a mutation... Then after this step is complete, I think that’s when embodied intelligence arrives." 3. He says that video models and world models have no relation to the intelligence ceiling, so DeepSeek isn't interested in them. 4. When he gave the speech (in May), DeepSeek had 20,000 H-equivalent GPUs: "We don’t have that many cards [GPUs]—the number of our cards is still relatively few. We currently have roughly 20,000 H-equivalent compute cards, and most of these just arrived, arrived in the last month or two, and there may be many more machines that haven’t arrived yet." Observation: OpenAI intends to run the automated AI research "intern" on the equivalent of 500,000 A-100 GPUs. This is something like ~8x the amount of compute that DeepSeek had available in May overall. 5. Export controls are working: "Our gap with the US is mainly in resources, and the gap in people isn’t very big... Talent isn’t the bottleneck—resources are the biggest bottleneck. Resources first affect talent cultivation, because with little compute, our opportunities to run experiments are relatively few, so our talent overall has a gap with the US. The talent gap is essentially also because of the compute gap. On the current largest models, we actually can’t afford to train. Even if we spent all 50 billion, we still couldn’t afford to train. Even if we could stack it up, we couldn’t afford to use it. If I wanted to train a model as large as [U.S. models], it should require 50,000 GB300s, or Huawei 950s, 200,000 cards. This is just training, not yet considering doing research. So the biggest gap between us and the US is in resources... This problem is currently basically unsolvable, because Huawei’s output is also limited. Because if I want to train 800B, I need 200,000 of Huawei’s newest cards, and this is just training, not yet considering doing research." 6. On Huawei and Chinese domestic chip production: "Huawei’s problem is still insufficient production capacity. Huawei gives us roughly 16,000 cards of capacity, internet giants maybe get a hundred-something thousand, we get ten-something thousand... this is probably just how much capacity Huawei has. So we also can’t count on training that bigger model on Huawei later, or training a model with several hundred B parameters active—there were some saying this year. But next year, the year after, there may be opportunities... ...the gap between [China] and the US on chips—I believe there won’t be a gap on ecosystem going forward, but on chips it’s fourfold plus two years... NVIDIA cards you can basically depreciate over five years. Huawei cards at most depreciate over three years. Huawei 950—this year using it is pretty good, next year using it I think is still okay, later using it I think might really be too power-hungry." 7. On the gap between China (or DeepSeek?) and the U.S.: "So the gap between us and the US might be lagging the US by 12 months, lagging the US maybe 12 to 18 months, or 6 to 12 months. Anyway, simply put, lagging the US by two years, then using only one-twentieth of the US’s compute to accomplish this thing. This narrative is lagging one to two years, but using only one-twentieth of its compute. So in the future we want to rewrite this narrative—that is, we use one-nth of its compute, but shorten this time even more, shorten it to 6 months, 3 months—I think this is a goal. And we can even surpass them in certain aspects. But under the situation where total compute still has an order-of-magnitude gap, comprehensive surpassing is unrealistic; but in certain key, trade-off-selected places, us surpassing in some areas might be possible."
I'm using this extremely helpful translated transcript from @GaoYingshi https://www.fredgao.com/p/deepseeks-liang-wenfeng-breaks-his
interesting aspect here is that when Liang talks about talent and DeepSeek being "just ordinary people, not geniuses", he doesn't mean IQ or gaokao or even IOI medals (many of his hires are in fact geniuses). He legit thinks that true talent *results* from hands-on "cultivation".
Key quotes from the Liang Wenfeng investor presentation: 1. He's convinced that current AI models can't reach AGI and we need continual learning: "Just like CoT—after CoT reached its ceiling, it already surpassed the most top-tier humans, in doing math olympiad problems and writing programs... But it still stops there, that technology couldn’t reach AGI. So you see, AI’s intelligence trajectory is traceable. [A]fter Agent, we think the problem to be solved should be continuous learning—that is, how to let the model learn continuously, rather than you giving it strong training—it should be able to do relatively long-term continuous learning like a human." Observation: This is very similar to how Google views AI progress and next steps (except that Liang does not care for world models and believes that robotics should come at a much later step; more on that below). 2. He thinks that the right order in which to tackle these problems is: continual learning, then RSI (which is a "gradual process", NOT an intelligence explosion), then embodied AI (robotics): "After continuous learning, we might arrive at a singularity. This singularity is: when this model can learn continuously, it can already do all the things humans can do. It could then develop its own version, could research on its own, then develop its own next version, developing more advanced AI models. So it would reach a singularity, able to achieve its own iteration. But this singularity—it’s not really a singularity, it’s also a gradual process. This process might also be a relatively long gradual change, not a mutation... Then after this step is complete, I think that’s when embodied intelligence arrives." 3. He says that video models and world models have no relation to the intelligence ceiling, so DeepSeek isn't interested in them. 4. When he gave the speech (in May), DeepSeek had 20,000 H-equivalent GPUs: "We don’t have that many cards [GPUs]—the number of our cards is still relatively few. We currently have roughly 20,000 H-equivalent compute cards, and most of these just arrived, arrived in the last month or two, and there may be many more machines that haven’t arrived yet." Observation: OpenAI intends to run the automated AI research "intern" on the equivalent of 500,000 A-100 GPUs. This is something like ~8x the amount of compute that DeepSeek had available in May overall. 5. Export controls are working: "Our gap with the US is mainly in resources, and the gap in people isn’t very big... Talent isn’t the bottleneck—resources are the biggest bottleneck. Resources first affect talent cultivation, because with little compute, our opportunities to run experiments are relatively few, so our talent overall has a gap with the US. The talent gap is essentially also because of the compute gap. On the current largest models, we actually can’t afford to train. Even if we spent all 50 billion, we still couldn’t afford to train. Even if we could stack it up, we couldn’t afford to use it. If I wanted to train a model as large as [U.S. models], it should require 50,000 GB300s, or Huawei 950s, 200,000 cards. This is just training, not yet considering doing research. So the biggest gap between us and the US is in resources... This problem is currently basically unsolvable, because Huawei’s output is also limited. Because if I want to train 800B, I need 200,000 of Huawei’s newest cards, and this is just training, not yet considering doing research." 6. On Huawei and Chinese domestic chip production: "Huawei’s problem is still insufficient production capacity. Huawei gives us roughly 16,000 cards of capacity, internet giants maybe get a hundred-something thousand, we get ten-something thousand... this is probably just how much capacity Huawei has. So we also can’t count on training that bigger model on Huawei later, or training a model with several hundred B parameters active—there were some saying this year. But next year, the year after, there may be opportunities... ...the gap between [China] and the US on chips—I believe there won’t be a gap on ecosystem going forward, but on chips it’s fourfold plus two years... NVIDIA cards you can basically depreciate over five years. Huawei cards at most depreciate over three years. Huawei 950—this year using it is pretty good, next year using it I think is still okay, later using it I think might really be too power-hungry." 7. On the gap between China (or DeepSeek?) and the U.S.: "So the gap between us and the US might be lagging the US by 12 months, lagging the US maybe 12 to 18 months, or 6 to 12 months. Anyway, simply put, lagging the US by two years, then using only one-twentieth of the US’s compute to accomplish this thing. This narrative is lagging one to two years, but using only one-twentieth of its compute. So in the future we want to rewrite this narrative—that is, we use one-nth of its compute, but shorten this time even more, shorten it to 6 months, 3 months—I think this is a goal. And we can even surpass them in certain aspects. But under the situation where total compute still has an order-of-magnitude gap, comprehensive surpassing is unrealistic; but in certain key, trade-off-selected places, us surpassing in some areas might be possible."
Combined views
6.6K
2 posts, first seen 5h ago