Announcement
EasyPPO is pitched as more stable PPO for LLM post-training
The EasyPPO announcement says it fixes the critic while leaving the actor update unchanged, with no new actor loss or policy algorithm.
TLDR
The author presents EasyPPO as a way to make PPO more stable during LLM post-training by fixing the critic rather than changing the actor update. They report better results and zero training collapse across all their experiments.
Combined views
15.9K
4 Sources, first seen 3h ago
172 likes
