Paper reports a distillation trade-off during mid-training: reasoning gains, slower factual learning
The authors propose Switch Distillation: use a teacher model’s predictions for tokens it is confident about, and fall back to cross-entropy training for the rest.
TLDR
The paper’s authors, quoted in a post, report that standard forward knowledge distillation—training with a teacher model’s predictions—behaves differently across training stages. With post-trained teachers, it improves reasoning and factual recall during pre-training relative to standard next-token prediction. During mid-training, they report continued reasoning gains but slower acquisition of factual recall.
Their proposed Switch Distillation uses the teacher’s predictive entropy to identify confident predictions, distilling on those tokens and otherwise using cross-entropy. A reply draws a parallel to research on down-scaling rather than distillation.
Combined views
523
1 Source, first seen 28d ago