• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researchers propose TailSFT to improve results after reinforcement learning

    A user’s summary says TailSFT records each training example’s loss before supervised fine-tuning, then excludes the most-improved examples in each batch from model updates.

    ND
    LB
    NL
    15 Sources, ,

    TLDR

    Researchers describe TailSFT as a lightweight way to improve coverage and performance after reinforcement learning (RL). A user summarizing the proposal says it first records each example’s loss before supervised fine-tuning, then excludes the examples that have improved most relative to that baseline from each batch’s gradient updates. That summary claims a better pass@k starting point for RL when k > 1.

    Combined views

    260.9K

    15 Sources, first seen 19d ago

    Combined views

    260.9K

    15 Sources, first seen 19d ago

    1.7K likes
    19d ago
    first seen 19d ago
    1.7K likes
    46 comments
    1.6K saves
    403 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    46 comments
    1.6K saves
    403 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    15 Sources

    @SadhikaMalladiRL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https://arxiv.org/abs/2510.15020). Now, we propose TailSFT (https://arxiv.org/abs/2608.25756), a lightweight + principled way to directly improve coverage and get better post-RL perf.
    @canondetortugasRT @SadhikaMalladi: RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https:…
    @gaotianyu1350It’s really cool to think about these LLM training stages end-to-end instead of optimizing each separately!
    @LianhuiqRT @SadhikaMalladi: RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https:…
    @giffmanaTL;DR: 1. Record pre-sft loss L0 of every example 2. During SFT, exclude the most-improved-compared-to-L0 examples in the batch from bprop (eg sort and slice or mul by zero) 3. Get a better pass@k for k>1 starting point for RL Seems simple and intuitive!
    @BlackHC@giffmana Also related RhoLoss, except this is the opposite in a way bc it uses the model under training as reference and flips the direction?
    @abeiramiSuper cool and simple method: - during training focus on training examples whose losses haven't dropped enough in each batch. The objective is intuitive similar to negatively tilting the samples with their initial reference, which is known to be equivalent to upweighting them.
    @natolambertCool post training work built on Olmo 3 :)
    @NandoDFRT @giffmana: TL;DR: 1. Record pre-sft loss L0 of every example 2. During SFT, exclude the most-improved-compared-to-L0 examples in the ba…
    @xiamengzhouRT @SadhikaMalladi: RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https:…

    15 Sources

    @SadhikaMalladiRL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https://arxiv.org/abs/2510.15020). Now, we propose TailSFT (https://arxiv.org/abs/2608.25756), a lightweight + principled way to directly improve coverage and get better post-RL perf.
    @canondetortugasRT @SadhikaMalladi: RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https:…
    @gaotianyu1350It’s really cool to think about these LLM training stages end-to-end instead of optimizing each separately!
    @LianhuiqRT @SadhikaMalladi: RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https:…
    @giffmanaTL;DR: 1. Record pre-sft loss L0 of every example 2. During SFT, exclude the most-improved-compared-to-L0 examples in the batch from bprop (eg sort and slice or mul by zero) 3. Get a better pass@k for k>1 starting point for RL Seems simple and intuitive!
    @BlackHC@giffmana Also related RhoLoss, except this is the opposite in a way bc it uses the model under training as reference and flips the direction?
    @abeiramiSuper cool and simple method: - during training focus on training examples whose losses haven't dropped enough in each batch. The objective is intuitive similar to negatively tilting the samples with their initial reference, which is known to be equivalent to upweighting them.
    @natolambertCool post training work built on Olmo 3 :)
    @NandoDFRT @giffmana: TL;DR: 1. Record pre-sft loss L0 of every example 2. During SFT, exclude the most-improved-compared-to-L0 examples in the ba…
    @xiamengzhouRT @SadhikaMalladi: RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https:…