Researchers propose TailSFT to improve results after reinforcement learning
A user’s summary says TailSFT records each training example’s loss before supervised fine-tuning, then excludes the most-improved examples in each batch from model updates.
TLDR
Researchers describe TailSFT as a lightweight way to improve coverage and performance after reinforcement learning (RL). A user summarizing the proposal says it first records each example’s loss before supervised fine-tuning, then excludes the examples that have improved most relative to that baseline from each batch’s gradient updates. That summary claims a better pass@k starting point for RL when k > 1.
Combined views
260.9K
15 Sources, first seen 19d ago