Language Compression Optimization Fuels Current AI
Researchers discuss Shannon's entropy estimates and their link to LLM capabilities.
Christian Szegedy posted that optimizing for compression of language produced the current level of AI. Peyman Milanfar wrote that Claude Shannon estimated English text is roughly 75 percent redundant, with long-range entropy at most 1 bit per letter. Milanfar added that text encoders reaching this rate appeared only recently and that modern LLMs now predict the next character at this level.
Combined views
81.6K
2 posts, first seen 8d ago
Language Compression Optimization Fuels Current AI
Researchers discuss Shannon's entropy estimates and their link to LLM capabilities.
Christian Szegedy posted that optimizing for compression of language produced the current level of AI. Peyman Milanfar wrote that Claude Shannon estimated English text is roughly 75 percent redundant, with long-range entropy at most 1 bit per letter. Milanfar added that text encoders reaching this rate appeared only recently and that modern LLMs now predict the next character at this level.