A proposed approach to OPD going off-policy after early mistakes in multi-turn tasks
A post proposes combining RL with reverse KL at the pivotal turn and forward KL on recovery rollouts from a privileged self-teacher.
TLDR
The poster says vanilla OPD often goes off-policy after early pivotal mistakes in multi-turn tasks. Their proposed approach combines reinforcement learning with reverse KL at the pivotal turn, then applies forward KL to recovery rollouts sampled from a privileged self-teacher.
Combined views
642
1 Source, first seen 2h ago
likes