• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    Training on a model’s own incorrect samples could beat distilling from a much larger teacher

    A user argues traditional RFT reinforces a narrow set of approaches, while diverse sampling yields better pass@k and test-time scaling.

    PM
    AG
    2 Sources, 8h ago,

    TLDR

    A user claims training a model on its own incorrect samples can beat distilling from a much larger teacher. They argue that strategically sampling for approach-level diversity yields better pass@k, test-time scaling and RL initialization than traditional RFT.

    Combined views

    1.3K

    2 Sources, first seen 8h ago

    Combined views

    1.3K

    2 Sources, first seen 8h ago

    31 likes
    first seen 8h ago
    31 likes
    2 comments
    27 saves
    8 reposts
    Featured Source
    2 comments
    27 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @AlexAag1234Did you know training on your own incorrect samples can beat distilling from a much larger teacher? Traditional RFT only reinforces a narrow set of approaches. Strategically sampling for approach-level diversity gives better pass@k, test-time scaling and RL initialization! https://alexgurung.me/diversity/ 🧵8h
    @PMinerviniRT @AlexAag1234: Did you know training on your own incorrect samples can beat distilling from a much larger teacher? Traditional RFT only…1h

    2 Sources

    @AlexAag1234Did you know training on your own incorrect samples can beat distilling from a much larger teacher? Traditional RFT only reinforces a narrow set of approaches. Strategically sampling for approach-level diversity gives better pass@k, test-time scaling and RL initialization! https://alexgurung.me/diversity/ 🧵8h
    @PMinerviniRT @AlexAag1234: Did you know training on your own incorrect samples can beat distilling from a much larger teacher? Traditional RFT only…1h