Tweet Proposes TailRL for Better RL Distributions
Post questions mean-focused RL and introduces TailRL to maximize coverage too.
TLDR
Shrinivas Ramasubramanian posted about whether reinforcement learning optimizes the right objective. He notes that standard methods focus on mean reward and often cause the distribution to collapse into a spike with the tail disappearing. The post introduces Tail-Likelihood Reinforcement Learning, called TailRL, which keeps mean reward maximization while also maximizing coverage. An attached video shows a reward density plot animation with on-screen text referencing a generative policy. The author lists an MS from Carnegie Mellon and undergrad studies at IIT Bombay.
Combined views
30.9K
2 Sources, first seen 28d ago