• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Paper reports a distillation trade-off during mid-training: reasoning gains, slower factual learning

    The authors propose Switch Distillation: use a teacher model’s predictions for tokens it is confident about, and fall back to cross-entropy training for the rest.

    DR
    1 Source, 28d ago, first seen 28d ago

    TLDR

    The paper’s authors, quoted in a post, report that standard forward knowledge distillation—training with a teacher model’s predictions—behaves differently across training stages. With post-trained teachers, it improves reasoning and factual recall during pre-training relative to standard next-token prediction. During mid-training, they report continued reasoning gains but slower acquisition of factual recall.

    Their proposed Switch Distillation uses the teacher’s predictive entropy to identify confident predictions, distilling on those tokens and otherwise using cross-entropy. A reply draws a parallel to research on down-scaling rather than distillation.

    Combined views

    523

    1 Source, first seen 28d ago

    Combined views

    523

    1 Source, first seen 28d ago

    5 likes
    5 likes
    1 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @roydanroy@iScienceLuvr Similar message to https://arxiv.org/abs/2310.04680, where they studied down-scaling, rather than distillation.

    1 Source

    @roydanroy@iScienceLuvr Similar message to https://arxiv.org/abs/2310.04680, where they studied down-scaling, rather than distillation.