Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Paper studies how forward KL distillation affects language model training at different stages.
TLDR
Tanishq Mathew Abraham posted on X about findings from an arXiv paper with the same title. The work examines logit-based knowledge distillation used to train smaller language models from stronger teachers. It reports that forward Kullback-Leibler distillation with post-trained teachers behaves differently during mid-training compared with other stages. The post includes quotes stating that forward KD simultaneously improves reasoning and factual recall under standard conditions. Code for the paper is hosted in a public GitHub repository maintained by facebookresearch.
Combined views
18.5K
4 Sources, first seen 29d ago