Shai Shalev-Shwartz Questions Best@k Gradient Estimation in RL
He highlights how current methods create diversity issues in expert agentic systems.
Shai Shalev-Shwartz, professor of computer science at Hebrew University of Jerusalem and CTO of Mobileye, posted thoughts on challenges in reinforcement learning for post-training. He pointed out that typical RL algorithms like GRPO optimize for top-1 performance. According to him this approach risks creating diversity problems. When building agentic systems for expert-level problems, the key measure becomes best@k instead. Shalev-Shwartz inquired about deriving an unbiased gradient estimator for best@k. He recommended related research appearing below and acknowledged the authors of that work.
Combined views
6.8K
1 post, first seen 15h ago
Shai Shalev-Shwartz Questions Best@k Gradient Estimation in RL
He highlights how current methods create diversity issues in expert agentic systems.
Shai Shalev-Shwartz, professor of computer science at Hebrew University of Jerusalem and CTO of Mobileye, posted thoughts on challenges in reinforcement learning for post-training. He pointed out that typical RL algorithms like GRPO optimize for top-1 performance. According to him this approach risks creating diversity problems. When building agentic systems for expert-level problems, the key measure becomes best@k instead. Shalev-Shwartz inquired about deriving an unbiased gradient estimator for best@k. He recommended related research appearing below and acknowledged the authors of that work.
