• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Standard SGD can reportedly match AdamW in LLM reinforcement learning

    A user summarizing research says AdamW's momentum and adaptive learning rates matter much less in reinforcement learning, questioning the need for its memory overhead.

    ET
    1 Source, 17d ago, first seen 17d ago

    TLDR

    A user summarizing research says standard stochastic gradient descent (SGD) works just as well as AdamW for reinforcement learning in large language models. The summary gives two reasons: most parameters need roughly the same effective step size, limiting the value of per-parameter adjustments, and accumulated momentum from past gradients is nearly unaligned with current gradients. As the user describes it, the authors argue that this training regime can drop AdamW entirely, avoiding its need to store extra gradient statistics in memory.

    Combined views

    14.8K

    1 Source, first seen 17d ago

    Combined views

    14.8K

    1 Source, first seen 17d ago

    243 likes
    243 likes
    15 comments
    213 saves
    24 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    15 comments
    213 saves
    24 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @ethantsliuYou don't need AdamW for RL in LLMs actually, standard SGD works just as well and is massively sparser! During RLVR, many default to using AdamW (which incurs massive memory overhead as it has to store both 1st and 2nd gradient moments). Existing assumptions strongly suggest that vanilla SGD performs poorly for large-scale transformers due to complex loss landscapes, making Adam's per-parameter adaptivity seem essential. The authors argue that RL is a fundamentally different training regime and that we can drop AdamW entirely. They show both momentum and adaptive learning rates are far less influential in RL, breaking down the reasoning into two distinct observations: 1) Adaptive learning rates aren't needed: In RL, the variance of the second moment (used for adaptive scaling) is ~22x lower than in SFT. This means most parameters require roughly the same effective step size, rendering per-parameter adaptivity largely redundant. 2) Momentum is counter-productive: RL landscapes are highly non-stationary because the data distribution co-evolves with the policy. The authors found that the cosine similarity between accumulated momentum and current gradients drops to near-zero in RL (unlike SFT where it is highly aligned), meaning past momentum points in unhelpful directions. To prove this, they stripped AdamW down to its components and tested vanilla SGD. Surprisingly, SGD naturally induces extreme parameter efficiency without any explicit regularization: Full fine-tuning with SGD updates fewer than 0.02% to 0.46% of model parameters. Because SGD lacks adaptive scaling, it avoids artificially amplifying irrelevant gradients. The memory savings are massive; dropping AdamW's momentum states saves 15.7 GB on just a 1.7B model! Tested across 3 domains (Math, Code, RLVE) and 2 algorithms (PPO, GRPO), SGD matches or beats AdamW while updating >1,000x fewer parameters!

    1 Source

    @ethantsliuYou don't need AdamW for RL in LLMs actually, standard SGD works just as well and is massively sparser! During RLVR, many default to using AdamW (which incurs massive memory overhead as it has to store both 1st and 2nd gradient moments). Existing assumptions strongly suggest that vanilla SGD performs poorly for large-scale transformers due to complex loss landscapes, making Adam's per-parameter adaptivity seem essential. The authors argue that RL is a fundamentally different training regime and that we can drop AdamW entirely. They show both momentum and adaptive learning rates are far less influential in RL, breaking down the reasoning into two distinct observations: 1) Adaptive learning rates aren't needed: In RL, the variance of the second moment (used for adaptive scaling) is ~22x lower than in SFT. This means most parameters require roughly the same effective step size, rendering per-parameter adaptivity largely redundant. 2) Momentum is counter-productive: RL landscapes are highly non-stationary because the data distribution co-evolves with the policy. The authors found that the cosine similarity between accumulated momentum and current gradients drops to near-zero in RL (unlike SFT where it is highly aligned), meaning past momentum points in unhelpful directions. To prove this, they stripped AdamW down to its components and tested vanilla SGD. Surprisingly, SGD naturally induces extreme parameter efficiency without any explicit regularization: Full fine-tuning with SGD updates fewer than 0.02% to 0.46% of model parameters. Because SGD lacks adaptive scaling, it avoids artificially amplifying irrelevant gradients. The memory savings are massive; dropping AdamW's momentum states saves 15.7 GB on just a 1.7B model! Tested across 3 domains (Math, Code, RLVE) and 2 algorithms (PPO, GRPO), SGD matches or beats AdamW while updating >1,000x fewer parameters!