Announcement
GRPO may not work as well with human-provided samples as with on-policy generated ones
A commenter says rewriting training data to optimize for surprise where it matters gives backpropagation a cleaner signal.
TLDR
A commenter reasons that rewriting training data to optimize for surprise only where it matters gives backpropagation a cleaner signal. They suggest GRPO might not work as well with human-provided samples instead of on-policy generated ones.
Combined views
58.1K
4 Sources, first seen ago