Paper author says reinforcement learning can teach language models new capabilities
The author says toy experiments also show that post-training with random rewards does not improve capabilities except under very particular circumstances.
TLDR
An author says the paper’s toy experiments show that reinforcement-learning post-training can teach language models capabilities not already present in the base model’s pass@k distribution—what it can get right across multiple attempts. The author also says random rewards do not improve capabilities except under very particular circumstances, and hopes the paper can serve as a tutorial on reinforcement-learning post-training.
Combined views
69.1K
3 Sources, first seen 19d ago