Schmidhuber Claims 1991 Invention of Neural Network Distillation
Reactions from ranked influencers
4 posts@pmddomingos distillation was published in 1991
In 2025, the DeepSeek “Sputnik" shocked the world, wiping out a trillion $ from the stock market. DeepSeek [7] distills knowledge from one neural network (NN) into another. Who invented this? https://people.idsia.ch/~juergen/who-invented-knowledge-distillation-with-neural-networks.html NN distillation was published in 1991 by yours truly [0]. Section 4 on a "conscious" chunker NN and a "subconscious” automatiser NN [0][1] introduced a general principle for transferring the knowledge of one NN to another. Suppose a teacher NN has learned to predict (conditional expectations of) data, given other data. Its knowledge can be compressed into a student NN, by training the student NN to imitate the behavior of the teacher NN (while also re-tarining the student NN on previously learned skills such that it does not forget them). In 1991, this was called "collapsing" or "compressing" the behavior of one NN into another. Today, this is widely used, and also referred to as “distilling" [2][6] or "cloning" the behavior of a teacher NN into that of a student NN. It even works when the NNs are recurrent and operate on different time scales [0][1]. See also [3][4]. REFERENCES (more in Technical Note IDSIA-12-25 [5]) [0] J. Schmidhuber. Neural sequence chunkers. Tech Report FKI-148-91, TU Munich, April 1991. [1] J. Schmidhuber. Learning complex, extended sequences using the principle of history compression. Neural Computation, 4(2):234-242, 1992. Based on [0]. [2] O. Vinyals, J. A. Dean, G. E. Hinton. Distilling the Knowledge in a Neural Network. Preprint arXiv:1503.02531 [http://stat.ML], 2015. The authors did not cite the original 1991 NN distillation procedure [0][1][DLP], not even in their later patent application. [3] J. Ba, R. Caruana. Do Deep Nets Really Need to be Deep? NIPS 2014. Preprint arXiv:1312.6184 (2013). [4] C. Bucilua, R. Caruana, and A. Niculescu-Mizil. Model compression. SIGKDD International conference on knowledge discovery and data mining, 2006. [5] J. Schmidhuber. Who invented knowledge distillation with artificial neural networks? Technical Note IDSIA-12-25, IDSIA, Nov 2025 [6] How 3 Turing awardees republished key methods and ideas whose creators they failed to credit. Technical Report IDSIA-23-23, 2023 [7] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Preprint arXiv:2501.12948, 2025
@mkratsios47 fun fact: distillation was published in 1991, then copied by others
In 2025, the DeepSeek “Sputnik" shocked the world, wiping out a trillion $ from the stock market. DeepSeek [7] distills knowledge from one neural network (NN) into another. Who invented this? https://people.idsia.ch/~juergen/who-invented-knowledge-distillation-with-neural-networks.html NN distillation was published in 1991 by yours truly [0]. Section 4 on a "conscious" chunker NN and a "subconscious” automatiser NN [0][1] introduced a general principle for transferring the knowledge of one NN to another. Suppose a teacher NN has learned to predict (conditional expectations of) data, given other data. Its knowledge can be compressed into a student NN, by training the student NN to imitate the behavior of the teacher NN (while also re-tarining the student NN on previously learned skills such that it does not forget them). In 1991, this was called "collapsing" or "compressing" the behavior of one NN into another. Today, this is widely used, and also referred to as “distilling" [2][6] or "cloning" the behavior of a teacher NN into that of a student NN. It even works when the NNs are recurrent and operate on different time scales [0][1]. See also [3][4]. REFERENCES (more in Technical Note IDSIA-12-25 [5]) [0] J. Schmidhuber. Neural sequence chunkers. Tech Report FKI-148-91, TU Munich, April 1991. [1] J. Schmidhuber. Learning complex, extended sequences using the principle of history compression. Neural Computation, 4(2):234-242, 1992. Based on [0]. [2] O. Vinyals, J. A. Dean, G. E. Hinton. Distilling the Knowledge in a Neural Network. Preprint arXiv:1503.02531 [http://stat.ML], 2015. The authors did not cite the original 1991 NN distillation procedure [0][1][DLP], not even in their later patent application. [3] J. Ba, R. Caruana. Do Deep Nets Really Need to be Deep? NIPS 2014. Preprint arXiv:1312.6184 (2013). [4] C. Bucilua, R. Caruana, and A. Niculescu-Mizil. Model compression. SIGKDD International conference on knowledge discovery and data mining, 2006. [5] J. Schmidhuber. Who invented knowledge distillation with artificial neural networks? Technical Note IDSIA-12-25, IDSIA, Nov 2025 [6] How 3 Turing awardees republished key methods and ideas whose creators they failed to credit. Technical Report IDSIA-23-23, 2023 [7] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Preprint arXiv:2501.12948, 2025
@Polymarket fun fact: distillation was published in 1991, then plagiarized by others
In 2025, the DeepSeek “Sputnik" shocked the world, wiping out a trillion $ from the stock market. DeepSeek [7] distills knowledge from one neural network (NN) into another. Who invented this? https://people.idsia.ch/~juergen/who-invented-knowledge-distillation-with-neural-networks.html NN distillation was published in 1991 by yours truly [0]. Section 4 on a "conscious" chunker NN and a "subconscious” automatiser NN [0][1] introduced a general principle for transferring the knowledge of one NN to another. Suppose a teacher NN has learned to predict (conditional expectations of) data, given other data. Its knowledge can be compressed into a student NN, by training the student NN to imitate the behavior of the teacher NN (while also re-tarining the student NN on previously learned skills such that it does not forget them). In 1991, this was called "collapsing" or "compressing" the behavior of one NN into another. Today, this is widely used, and also referred to as “distilling" [2][6] or "cloning" the behavior of a teacher NN into that of a student NN. It even works when the NNs are recurrent and operate on different time scales [0][1]. See also [3][4]. REFERENCES (more in Technical Note IDSIA-12-25 [5]) [0] J. Schmidhuber. Neural sequence chunkers. Tech Report FKI-148-91, TU Munich, April 1991. [1] J. Schmidhuber. Learning complex, extended sequences using the principle of history compression. Neural Computation, 4(2):234-242, 1992. Based on [0]. [2] O. Vinyals, J. A. Dean, G. E. Hinton. Distilling the Knowledge in a Neural Network. Preprint arXiv:1503.02531 [http://stat.ML], 2015. The authors did not cite the original 1991 NN distillation procedure [0][1][DLP], not even in their later patent application. [3] J. Ba, R. Caruana. Do Deep Nets Really Need to be Deep? NIPS 2014. Preprint arXiv:1312.6184 (2013). [4] C. Bucilua, R. Caruana, and A. Niculescu-Mizil. Model compression. SIGKDD International conference on knowledge discovery and data mining, 2006. [5] J. Schmidhuber. Who invented knowledge distillation with artificial neural networks? Technical Note IDSIA-12-25, IDSIA, Nov 2025 [6] How 3 Turing awardees republished key methods and ideas whose creators they failed to credit. Technical Report IDSIA-23-23, 2023 [7] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Preprint arXiv:2501.12948, 2025
fun fact: distillation was published in 1991 and subsequently copied by others without citing the source https://x.com/SchmidhuberAI/status/1988992765141614831
We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model. To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection. Moonshot AI has also acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models. The United States strongly supports the free and fair development of AI, including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models. Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.
Combined views
5.6K
4 posts, first seen 5h ago