• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Tweet Proposes TailRL for Better RL Distributions

    Post questions mean-focused RL and introduces TailRL to maximize coverage too.

    GS
    SR
    2 Sources, 28d ago, first seen 28d ago

    TLDR

    Shrinivas Ramasubramanian posted about whether reinforcement learning optimizes the right objective. He notes that standard methods focus on mean reward and often cause the distribution to collapse into a spike with the tail disappearing. The post introduces Tail-Likelihood Reinforcement Learning, called TailRL, which keeps mean reward maximization while also maximizing coverage. An attached video shows a reward density plot animation with on-screen text referencing a generative policy. The author lists an MS from Carnegie Mellon and undergrad studies at IIT Bombay.

    Combined views

    30.9K

    2 Sources, first seen 28d ago

    Combined views

    30.9K

    2 Sources, first seen 28d ago

    310 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    310 likes
    5 comments
    319 saves
    85 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 comments
    319 saves
    85 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @stablegradientsIs RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
    @g_k_swamyRT @stablegradients: Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the…

    2 Sources

    @stablegradientsIs RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
    @g_k_swamyRT @stablegradients: Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the…