Announcement
A 1991 paper's claimed role in three techniques used by modern language models
Jürgen Schmidhuber says his paper introduced deep-net pre-training, positional encoding and neural-network distillation.
TLDR
Jürgen Schmidhuber says his 1991 paper, “Neural sequence chunkers,” introduced three techniques he considers essential to modern large language models: pre-training for deep neural networks, positional encoding and distilling knowledge from one neural network to another. His overview describes networks trained to predict their next input, passing unexpected inputs to higher levels.
Combined views
10K
2 Sources, first seen ago
