A post breaks down importance sampling in reinforcement learning
The explainer describes a training mismatch: samples can come from a policy that differs from the one being optimized, including when algorithms reuse the same samples for multiple updates.
TLDR
Importance sampling can help correct mismatches between the policy generating training samples and the policy being optimized, the post explains. The technique weights samples using the ratio of their probability under a target distribution to their probability under the distribution they came from. Samples more likely under the target distribution receive larger weights. The post highlights PPO, a reinforcement-learning algorithm that uses this ratio to compare a sampled token’s probability under the current policy with its probability under the policy that generated it.
Combined views
679
1 Source, first seen 20d ago