• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Paper author says reinforcement learning can teach language models new capabilities

    The author says toy experiments also show that post-training with random rewards does not improve capabilities except under very particular circumstances.

    NJ
    SH
    SH
    3 Sources, ,

    TLDR

    An author says the paper’s toy experiments show that reinforcement-learning post-training can teach language models capabilities not already present in the base model’s pass@k distribution—what it can get right across multiple attempts. The author also says random rewards do not improve capabilities except under very particular circumstances, and hopes the paper can serve as a tutorial on reinforcement-learning post-training.

    Combined views

    69.1K

    3 Sources, first seen 19d ago

    Combined views

    69.1K

    3 Sources, first seen 19d ago

    470 likes
    19d ago
    first seen 19d ago
    470 likes
    9 comments
    496 saves
    44 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    9 comments
    496 saves
    44 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @sankar_harilalCan RL post-training improve model capabilities with random rewards? And can it teach skills outside the base model's distribution? Our new paper, Demystifying Reinforcement Learning Post-training of Language Models, unpacks why it succeeds or fails. https://arxiv.org/abs/2608.24949 🧵
    @natashajaquesNo, RL post-training on random rewards does not improve model capabilities, except under very particular circumstances. Yes, RL post-training can teach models capabilities that aren’t already present in the base model’s pass@k distribution. While these findings might be obvious to RL folks, in this paper we conduct careful toy experiments to show they are also true for LLM post-training. Hoping this paper can be a useful tutorial for folks looking to learn more about RL post-training of LLMs.
    @ShikharMurtyRT @sankar_harilal: Can RL post-training improve model capabilities with random rewards? And can it teach skills outside the base model's d…

    3 Sources

    @sankar_harilalCan RL post-training improve model capabilities with random rewards? And can it teach skills outside the base model's distribution? Our new paper, Demystifying Reinforcement Learning Post-training of Language Models, unpacks why it succeeds or fails. https://arxiv.org/abs/2608.24949 🧵
    @natashajaquesNo, RL post-training on random rewards does not improve model capabilities, except under very particular circumstances. Yes, RL post-training can teach models capabilities that aren’t already present in the base model’s pass@k distribution. While these findings might be obvious to RL folks, in this paper we conduct careful toy experiments to show they are also true for LLM post-training. Hoping this paper can be a useful tutorial for folks looking to learn more about RL post-training of LLMs.
    @ShikharMurtyRT @sankar_harilal: Can RL post-training improve model capabilities with random rewards? And can it teach skills outside the base model's d…