Report
Training on a model’s own incorrect samples could beat distilling from a much larger teacher
A user argues traditional RFT reinforces a narrow set of approaches, while diverse sampling yields better pass@k and test-time scaling.
TLDR
A user claims training a model on its own incorrect samples can beat distilling from a much larger teacher. They argue that strategically sampling for approach-level diversity yields better pass@k, test-time scaling and RL initialization than traditional RFT.
Combined views
1.3K
2 Sources, first seen ago
