• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Shai Shalev-Shwartz Questions Best@k Gradient Estimation in RL

    He highlights how current methods create diversity issues in expert agentic systems.

    FP
    SS
    2 Sources, 28d ago, first seen 28d ago

    TLDR

    Shai Shalev-Shwartz, professor of computer science at Hebrew University of Jerusalem and CTO of Mobileye, posted thoughts on challenges in reinforcement learning for post-training. He pointed out that typical RL algorithms like GRPO optimize for top-1 performance. According to him this approach risks creating diversity problems. When building agentic systems for expert-level problems, the key measure becomes best@k instead. Shalev-Shwartz inquired about deriving an unbiased gradient estimator for best@k. He recommended related research appearing below and acknowledged the authors of that work.

    Combined views

    7.4K

    2 Sources, first seen 28d ago

    Combined views

    7.4K

    2 Sources, first seen 28d ago

    125 likes
    125 likes
    6 comments
    124 saves
    19 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 comments
    124 saves
    19 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @shai_s_shwartzRL algorithms for post-training (e.g., GRPO) typically optimize top-1 performance. But this can create a diversity problem. In agentic systems tackling expert-level problems, what we often really care about is best@k. So how do you derive an unbiased gradient estimator for best@k? See the nice work below 👇 Kudos to my academic grandson @NadavSchweiger! https://www.doubleai.com/research/argmaxrl-generalizing-maxrl-to-continuous-rewards
    @fpedregosaRT @shai_s_shwartz: RL algorithms for post-training (e.g., GRPO) typically optimize top-1 performance. But this can create a diversity prob…

    2 Sources

    @shai_s_shwartzRL algorithms for post-training (e.g., GRPO) typically optimize top-1 performance. But this can create a diversity problem. In agentic systems tackling expert-level problems, what we often really care about is best@k. So how do you derive an unbiased gradient estimator for best@k? See the nice work below 👇 Kudos to my academic grandson @NadavSchweiger! https://www.doubleai.com/research/argmaxrl-generalizing-maxrl-to-continuous-rewards
    @fpedregosaRT @shai_s_shwartz: RL algorithms for post-training (e.g., GRPO) typically optimize top-1 performance. But this can create a diversity prob…