Switch Distillation targets the tradeoff between reasoning and factual recall
The researchers introducing Switch Distillation say it uses a teacher model’s uncertainty to improve the tradeoff between reasoning and factual recall.
TLDR
The researchers report that knowledge distillation behaves differently across training stages. Compared with standard next-token prediction, they say it improves reasoning and factual recall in pre-training, but trades off recall for reasoning in mid-training. They introduce Switch Distillation, which they say improves that tradeoff using the teacher model’s uncertainty.
Combined views
7.4K
1 Source, first seen 15d ago