Modified model looping could yield compute-efficiency gains that grow with scale
A paper coauthor says modifying recursive depth, or looping, for model growth can improve pre-training scaling exponents.
TLDR
Scaling laws predict how loss decreases as computation increases. The paper’s authors say their modification to recursive depth—looping—for model growth can improve pre-training scaling exponents, meaning compute-efficiency gains that increase with scale.
Combined views
22.8K
2 Sources, first seen ago