Length-weighted reward baseline reportedly made rollout–trainer numerics more stable
The post describes replacing GRPO’s group mean reward baseline with a length-weighted mean, giving longer rollouts more influence on the reference value.
TLDR
A post describes moving from GRPO’s group mean reward baseline to a length-weighted group baseline. Each rollout’s reward Rᵢ is weighted by its length Lᵢ, and the resulting baseline is subtracted from the reward: R − (Σ RᵢLᵢ)/(Σ Lᵢ). The author says the change made rollout–trainer numerics more stable.
Combined views
14.6K
3 Sources, first seen 19d ago