Salakhutdinov shares TailRL work on arXiv
Russ Salakhutdinov posted about new work extending maximum-likelihood reinforcement learning to continuous rewards.
TLDR
Russ Salakhutdinov announced Tail-Likelihood Reinforcement Learning, or TailRL. The method extends maximum-likelihood reinforcement learning from binary to continuous rewards. Instead of optimizing only mean reward, TailRL maximizes the expected log of upper-tail probabilities to place more weight on rare high-reward outcomes. The arXiv post notes that two policies can share the same average reward yet differ sharply in their tail behavior. The paper is available at the linked source.
Combined views
47K
1 Source, first seen 24d ago