• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

    Paper studies how forward KL distillation affects language model training at different stages.

    TM
    SL
    AL
    4 Sources, ,

    TLDR

    Tanishq Mathew Abraham posted on X about findings from an arXiv paper with the same title. The work examines logit-based knowledge distillation used to train smaller language models from stronger teachers. It reports that forward Kullback-Leibler distillation with post-trained teachers behaves differently during mid-training compared with other stages. The post includes quotes stating that forward KD simultaneously improves reasoning and factual recall under standard conditions. Code for the paper is hosted in a public GitHub repository maintained by facebookresearch.

    Combined views

    18.5K

    4 Sources, first seen 29d ago

    Combined views

    18.5K

    4 Sources, first seen 29d ago

    335 likes
    29d ago
    first seen 29d ago
    335 likes
    8 comments
    259 saves
    66 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    8 comments
    259 saves
    66 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    @iScienceLuvrKnowledge Distillation During Mid-Training Favors Reasoning over Factual Recall "we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training" "while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains." "we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy" code: https://github.com/facebookresearch/midtraining-distillation paper link: https://arxiv.org/abs/2609.01532
    @StellaLisyRT @iScienceLuvr: Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall "we find that forward Kullback-Leibler (…
    @askalphaxiv“Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall” Knowledge distillation helps reasoning during mid-training, but surprisingly hurts factual recall because teachers are much less confident on the remaining knowledge-heavy tokens. So this paper introduces Switch Distillation, which distills only when teacher entropy is low and falls back to next-token prediction otherwise. It gets up to 1.71x better reasoning and 1.19x better knowledge and commonsense while preserving ~97% of factual recall. https://www.alphaxiv.org/abs/2609.01532

    4 Sources

    @iScienceLuvrKnowledge Distillation During Mid-Training Favors Reasoning over Factual Recall "we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training" "while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains." "we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy" code: https://github.com/facebookresearch/midtraining-distillation paper link: https://arxiv.org/abs/2609.01532
    @StellaLisyRT @iScienceLuvr: Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall "we find that forward Kullback-Leibler (…
    @askalphaxiv“Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall” Knowledge distillation helps reasoning during mid-training, but surprisingly hurts factual recall because teachers are much less confident on the remaining knowledge-heavy tokens. So this paper introduces Switch Distillation, which distills only when teacher entropy is low and falls back to next-token prediction otherwise. It gets up to 1.71x better reasoning and 1.19x better knowledge and commonsense while preserving ~97% of factual recall. https://www.alphaxiv.org/abs/2609.01532