• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Microsoft UCSD Paper Claims SFT Hurts RL Starting Points

    Tweet notes TailSFT preserves rare correct behaviors better for reinforcement learning.

    RP
    2 Sources, 27d ago, first seen 27d ago

    TLDR

    Rohan Paul posted that a Microsoft and University of San Diego paper shows supervised fine-tuning can make a model look stronger while leaving it weaker as a base for reinforcement learning. The post states SFT may remove rare correct behaviors that RL needs to discover. It adds that TailSFT keeps more of those behaviors and, under the same RL setup, reached up to 3.93 percentage points higher. The tweet includes a screenshot of one academic figure.

    Combined views

    6.2K

    2 Sources, first seen 27d ago

    Combined views

    6.2K

    2 Sources, first seen 27d ago

    36 likes
    36 likes
    3 comments
    19 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    19 saves
    11 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @rohanpaul_aiNew Microsoft plus Univ of San Diego paper shows a model can look better after SFT but actually be a worse starting point for RL, because SFT may wipe out rare correct behaviors that RL needs to discover. TailSFT preserves more of those behaviors, and with the same RL setup it produced up to 3.93 percentage points higher final pass@1. standard SFT can make RL harder by overtraining already-fit examples; TailSFT filters them and gives the later RL stage a better starting point. TailSFT changes SFT by filtering sequences whose loss has already dropped most relative to the base model. That shifts training toward examples the model still underfits, with one goal: keep correct responses reachable under repeated sampling so RL has useful behavior to reinforce.

    2 Sources

    @rohanpaul_aiNew Microsoft plus Univ of San Diego paper shows a model can look better after SFT but actually be a worse starting point for RL, because SFT may wipe out rare correct behaviors that RL needs to discover. TailSFT preserves more of those behaviors, and with the same RL setup it produced up to 3.93 percentage points higher final pass@1. standard SFT can make RL harder by overtraining already-fit examples; TailSFT filters them and gives the later RL stage a better starting point. TailSFT changes SFT by filtering sequences whose loss has already dropped most relative to the base model. That shifts training toward examples the model still underfits, with one goal: keep correct responses reachable under repeated sampling so RL has useful behavior to reinforce.