Switch Distillation aims to improve the reasoning–recall tradeoff in AI training
The researchers introducing the method say it uses the teacher model’s uncertainty to improve a tradeoff they found during mid-training: better reasoning at the expense of factual recall.
TLDR
The team behind Switch Distillation reports that knowledge distillation behaves differently across training stages. Compared with standard next-token prediction, it improves reasoning and factual recall in pre-training but trades off recall for reasoning in mid-training, the team says. The researchers introduce Switch Distillation as a method that uses teacher uncertainty to improve that tradeoff.
Combined views
2
1 Source, first seen 15d ago