Testing the small-model limits of AI scaling laws
A reader of a FAIR and NYU paper says models below perhaps 64 million parameters need intensive tuning—and warns against equating easier tuning at larger tested scales with simpler training.
TLDR
A reader describes a paper by FAIR and NYU researchers exploring how far scaling laws extend to small models without losing fit or predictability. According to the thread, the researchers found that good training settings became easier to find as scale increased within their tests. The reader cautions that this does not mean training gets simpler: additional instabilities arise beyond the study’s scale range. They suggest models below perhaps 64 million parameters need intensive tuning. The thread also describes the paper’s finding that different ways of counting computation and parameters matter much less than tuning, while cautioning that small discrepancies in the low-loss region can still be significant.
Testing the small-model limits of AI scaling laws
A reader of a FAIR and NYU paper says models below perhaps 64 million parameters need intensive tuning—and warns against equating easier tuning at larger tested scales with simpler training.
TLDR
A reader describes a paper by FAIR and NYU researchers exploring how far scaling laws extend to small models without losing fit or predictability. According to the thread, the researchers found that good training settings became easier to find as scale increased within their tests. The reader cautions that this does not mean training gets simpler: additional instabilities arise beyond the study’s scale range. They suggest models below perhaps 64 million parameters need intensive tuning. The thread also describes the paper’s finding that different ways of counting computation and parameters matter much less than tuning, while cautioning that small discrepancies in the low-loss region can still be significant.
