LLM token-loss gaps reportedly stay roughly constant across optimizers but not changes to training data
A researcher sharing a paper calls this pattern relative generalization invariance: two models’ token-wise losses differ by a constant gap. They say it largely holds across optimizers and moderate architecture changes, but changing the training data stream breaks it.
TLDR
A researcher sharing a paper says two models satisfy relative generalization invariance when their token-wise losses differ by a constant gap. They report that this pattern largely holds across optimizers and moderate architecture changes, producing an approximately uniform shift in token-wise loss. Changing the training data stream breaks the pattern, they say.
Combined views
54
1 Source, first seen 6h ago