• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    GRPO may not work as well with human-provided samples as with on-policy generated ones

    A commenter says rewriting training data to optimize for surprise where it matters gives backpropagation a cleaner signal.

    EG
    DH
    M🍥
    4 Sources, ,

    TLDR

    A commenter reasons that rewriting training data to optimize for surprise only where it matters gives backpropagation a cleaner signal. They suggest GRPO might not work as well with human-provided samples instead of on-policy generated ones.

    Combined views

    58.1K

    4 Sources, first seen 5h ago

    Combined views

    58.1K

    4 Sources, first seen 5h ago

    729 likes
    5h ago
    first seen 5h ago
    729 likes
    20 comments
    746 saves
    104 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    20 comments
    746 saves
    104 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #9

    Today's Rank

    #9

    4 Sources

    @aakaran31SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. 🤯 Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack. 1/n5h
    @mayferintuitively makes sense, by rewriting training data to optimize for surprise only where it matters the backprop is given much cleaner signal. with this line of thinking one could expect GRPO to not work that well when using N human provided samples instead of on policy generated4h
    @dhadfieldmenellRT @aakaran31: SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and…4h
    @egrefenVery cool work2h

    4 Sources

    @aakaran31SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. 🤯 Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack. 1/n5h
    @mayferintuitively makes sense, by rewriting training data to optimize for surprise only where it matters the backprop is given much cleaner signal. with this line of thinking one could expect GRPO to not work that well when using N human provided samples instead of on policy generated4h
    @dhadfieldmenellRT @aakaran31: SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and…4h
    @egrefenVery cool work2h