TailRL researchers say their method gives rare, high-reward runs more weight
The researchers say TailRL extends maximum-likelihood reinforcement learning from binary to continuous rewards and needs only a simple change to fit into existing training pipelines.
TLDR
The team behind Tail-Likelihood Reinforcement Learning (TailRL) says its method goes beyond optimizing only mean reward, giving rare, high-reward runs more weight. It does this by optimizing an objective based on the probability of rewards at the upper end of the range. The researchers report that across object localization, maze navigation, GUI grounding and code optimization, TailRL effectively uses rare high-reward samples and scales better with increased inference-time sampling.
Combined views
1 Source, first seen 23d ago