A post explains importance sampling’s role in reinforcement learning
The explainer describes how probability ratios can help correct a mismatch between the policy that generated training examples and the policy being optimized.
TLDR
The post describes importance sampling as a way to estimate an expectation under one probability distribution using samples drawn from another. It weights each sample by the ratio of its probability under the target distribution to its probability under the sampling distribution. The explainer connects this to PPO, a reinforcement-learning algorithm that may update a policy multiple times using the same generated examples. It says a core part of PPO’s loss function compares a sampled token’s probability under the current policy with its probability under the policy that generated it.
Combined views
7.7K
1 Source, first seen 19d ago