Standard SGD can reportedly match AdamW in LLM reinforcement learning
A user summarizing research says AdamW's momentum and adaptive learning rates matter much less in reinforcement learning, questioning the need for its memory overhead.
TLDR
A user summarizing research says standard stochastic gradient descent (SGD) works just as well as AdamW for reinforcement learning in large language models. The summary gives two reasons: most parameters need roughly the same effective step size, limiting the value of per-parameter adjustments, and accumulated momentum from past gradients is nearly unaligned with current gradients. As the user describes it, the authors argue that this training regime can drop AdamW entirely, avoiding its need to store extra gradient statistics in memory.
Combined views
14.8K
1 Source, first seen 17d ago