Byte-level models reportedly pull ahead as compute grows in Meta research
A post discussing a Meta paper says its distilled 1-billion-parameter byte models matched a distilled token model with one-sixth of the training data.
TLDR
A post about a Meta paper describes distilled 1-billion-parameter models trained on up to 1 trillion bytes. It says token models lead at low compute but plateau, while byte-level models overtake them as compute grows. According to the post, the work trains byte-level students by converting token-based teachers’ output scores into byte-level scores. Its exact conversion method is called End-Of-Token. The post says fitted scaling laws predict that model will end up to 4% ahead of the distilled token model. It also reports that a 256-entry vocabulary cuts storage for teacher output scores to about a fifth.
Combined views
54.6K
3 Sources, first seen 16d ago