Reaction
Randomized training spans may improve full-sequence loss faster in a toy test
A user says varying power-of-two spans from 128 to 512 for 20% of samples showed apparent gains.
TLDR
A user argues that longer sequences improve future-token prediction on average by providing more conditional information, while in-context learning skews improvements toward later tokens. In a toy baseline, they say randomizing power-of-two spans from 128 to 512 for 20% of samples seemed to improve full-sequence loss faster, despite less total information being available.
Combined views
2.4K
4 Sources, first seen ago