• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Length-weighted reward baseline reportedly made rollout–trainer numerics more stable

    The post describes replacing GRPO’s group mean reward baseline with a length-weighted mean, giving longer rollouts more influence on the reference value.

    KK
    YJ
    SH
    3 Sources, ,

    TLDR

    A post describes moving from GRPO’s group mean reward baseline to a length-weighted group baseline. Each rollout’s reward Rᵢ is weighted by its length Lᵢ, and the resulting baseline is subtracted from the reward: R − (Σ RᵢLᵢ)/(Σ Lᵢ). The author says the change made rollout–trainer numerics more stable.

    Combined views

    14.6K

    3 Sources, first seen 19d ago

    Combined views

    14.6K

    3 Sources, first seen 19d ago

    153 likes
    19d ago
    first seen 19d ago
    153 likes
    174 saves
    23 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    174 saves
    23 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @YouJiachengRT @RichardYRLi: Nice to see a KKT (duality) argument in modern LLM RL applications.
    @shikhargupta02They also moved away from using group mean reward (from GRPO) as variance reduction baseline to a length weighted group baseline - weight each rollout by its length Lᵢ which gave them more stable rollout <-> trainer numerics - (R - (Σ RᵢLᵢ) / (Σ Lᵢ))
    @kastnerkyleRT @shikhargupta02: They also moved away from using group mean reward (from GRPO) as variance reduction baseline to a length weighted group…

    3 Sources

    @YouJiachengRT @RichardYRLi: Nice to see a KKT (duality) argument in modern LLM RL applications.
    @shikhargupta02They also moved away from using group mean reward (from GRPO) as variance reduction baseline to a length weighted group baseline - weight each rollout by its length Lᵢ which gave them more stable rollout <-> trainer numerics - (R - (Σ RᵢLᵢ) / (Σ Lᵢ))
    @kastnerkyleRT @shikhargupta02: They also moved away from using group mean reward (from GRPO) as variance reduction baseline to a length weighted group…