Microsoft UCSD Paper Claims SFT Hurts RL Starting Points
Tweet notes TailSFT preserves rare correct behaviors better for reinforcement learning.
TLDR
Rohan Paul posted that a Microsoft and University of San Diego paper shows supervised fine-tuning can make a model look stronger while leaving it weaker as a base for reinforcement learning. The post states SFT may remove rare correct behaviors that RL needs to discover. It adds that TailSFT keeps more of those behaviors and, under the same RL setup, reached up to 3.93 percentage points higher. The tweet includes a screenshot of one academic figure.
Combined views
6.2K
2 Sources, first seen 27d ago