• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    A post explains importance sampling’s role in reinforcement learning

    The explainer describes how probability ratios can help correct a mismatch between the policy that generated training examples and the policy being optimized.

    CR
    1 Source, 19d ago, first seen 19d ago

    TLDR

    The post describes importance sampling as a way to estimate an expectation under one probability distribution using samples drawn from another. It weights each sample by the ratio of its probability under the target distribution to its probability under the sampling distribution. The explainer connects this to PPO, a reinforcement-learning algorithm that may update a policy multiple times using the same generated examples. It says a core part of PPO’s loss function compares a sampled token’s probability under the current policy with its probability under the policy that generated it.

    Combined views

    7.7K

    1 Source, first seen 19d ago

    Combined views

    7.7K

    1 Source, first seen 19d ago

    177 likes
    177 likes
    6 comments
    191 saves
    23 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 comments
    191 saves
    23 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @cwolferesearchImportance sampling is a concept that appears constantly in RL research (e.g., PPO objective, training / inference mismatch, and more). Here’s how it works… In RL, the policy used to generate rollouts does not always perfectly match the current policy we are optimizing. For example, algorithms like PPO may perform multiple policy updates over the same rollouts, while recent asynchronous RL infrastructure may lead to the incorporation of mildly stale or off-policy data into the training process. Importance sampling is a general technique in probability theory that can help to correct mismatches of this kind. Formally, importance sampling allows us to estimate an expectation under a target distribution f(x) using samples drawn from a different proposal distribution g(x). Instead of sampling directly from f(x), we can sample from g(x) and correct for the discrepancy between these distributions using the importance ratio f(x) / g(x). Intuitively, samples that are more likely under f(x) than g(x) receive a larger importance ratio, while samples that are less likely receive a smaller ratio. Importance ratios appear constantly in research on RL for LLMs. For example, PPO uses an importance ratio to compare the probability of a sampled token under the current policy and the policy that sampled the rollout—this ratio is a core component of the PPO loss function. Additionally, if rollouts are generated using a policy that is different than the policy being optimized—or a separate inference engine that produces slightly different token distributions—we can use importance ratios to account for this mismatch. In practice, the importance ratio can become large when the distributions differ substantially, yielding unstable and high-variance estimates. For this reason, RL algorithms may choose to clip or truncate the importance ratio, which introduces bias in order to reduce variance and improve stability. For example, both PPO and truncated importance sampling (TIS) adopt a clipped version of the importance ratio.

    1 Source

    @cwolferesearchImportance sampling is a concept that appears constantly in RL research (e.g., PPO objective, training / inference mismatch, and more). Here’s how it works… In RL, the policy used to generate rollouts does not always perfectly match the current policy we are optimizing. For example, algorithms like PPO may perform multiple policy updates over the same rollouts, while recent asynchronous RL infrastructure may lead to the incorporation of mildly stale or off-policy data into the training process. Importance sampling is a general technique in probability theory that can help to correct mismatches of this kind. Formally, importance sampling allows us to estimate an expectation under a target distribution f(x) using samples drawn from a different proposal distribution g(x). Instead of sampling directly from f(x), we can sample from g(x) and correct for the discrepancy between these distributions using the importance ratio f(x) / g(x). Intuitively, samples that are more likely under f(x) than g(x) receive a larger importance ratio, while samples that are less likely receive a smaller ratio. Importance ratios appear constantly in research on RL for LLMs. For example, PPO uses an importance ratio to compare the probability of a sampled token under the current policy and the policy that sampled the rollout—this ratio is a core component of the PPO loss function. Additionally, if rollouts are generated using a policy that is different than the policy being optimized—or a separate inference engine that produces slightly different token distributions—we can use importance ratios to account for this mismatch. In practice, the importance ratio can become large when the distributions differ substantially, yielding unstable and high-variance estimates. For this reason, RL algorithms may choose to clip or truncate the importance ratio, which introduces bias in order to reduce variance and improve stability. For example, both PPO and truncated importance sampling (TIS) adopt a clipped version of the importance ratio.