Sharpness-Aware Multiscale Momentum (EMA-Mix)
A user calls the approach “quite reasonable,” sharing an arXiv paper whose description highlights the computational cost of LLM pretraining.
TLDR
A user shares an arXiv paper with a favorable assessment of Sharpness-Aware Multiscale Momentum (EMA-Mix). The paper’s description says pretraining accounts for a large fraction of LLM training’s computational cost. It identifies noise-dominant gradients and a highly ill-conditioned loss landscape as major optimization challenges.
Combined views
20.9K
2 Sources, first seen 3d ago
Sharpness-Aware Multiscale Momentum (EMA-Mix)
A user calls the approach “quite reasonable,” sharing an arXiv paper whose description highlights the computational cost of LLM pretraining.
TLDR
A user shares an arXiv paper with a favorable assessment of Sharpness-Aware Multiscale Momentum (EMA-Mix). The paper’s description says pretraining accounts for a large fraction of LLM training’s computational cost. It identifies noise-dominant gradients and a highly ill-conditioned loss landscape as major optimization challenges.