Report
Could privileged information help reinforcement learning find useful rollouts?
A post asks whether privileged information could guide which rollouts to sample, rather than only score them afterward.
TLDR
A post introducing “Pedagogical RL” argues that typical reinforcement-learning algorithms and on-policy distillation use privileged information to score rollouts, but not to find them. It asks whether that information could instead help sample rollouts that reinforcement learning might otherwise stumble upon through more computation.
Combined views
92
1 Source, first seen ago
