Compression Equals Next-Word Prediction in Language Models
Reactions from ranked influencers
2 postsCompression *is* next-word prediction. How well you compress depends on how well you (learn to) predict. A universal compressor reaches the entropy rate of the underlying data process.
As someone who's been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is the closest thing to an ML PhD qualifying exam I've ever seen released publicly for free. Everyone thinks language models predict the next word. They don't. They compress language. Once you see the math, you can't unsee it. 33 minutes. Bookmark & watch today.
Compression *is* next-word prediction. How well you compress depends on how well you (learn to) predict the next word. A universal compressor learns the underlying data process and asymptotically compresses data to its entropy rate, which is the fundamental compression limit.
As someone who's been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is the closest thing to an ML PhD qualifying exam I've ever seen released publicly for free. Everyone thinks language models predict the next word. They don't. They compress language. Once you see the math, you can't unsee it. 33 minutes. Bookmark & watch today.
Combined views
1.7K
2 posts, first seen 3h ago