• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Cameron Wolfe Questions PPO Ratio Derivation in Tweet

    Netflix scientist shares thoughts on common PPO explanation gaps.

    CR
    1 Source, 29d ago, first seen 29d ago

    TLDR

    Cameron R. Wolfe, Staff Research Scientist at Netflix and author of the Deep Learning Focus Substack, posted on X about the policy importance ratio in PPO. He observes that typical explainers begin with a vanilla policy gradient expression and then move directly to PPO. The post points out that PPO is usually presented as a loss multiplying an importance ratio by the advantage. Wolfe asks where that ratio originates and indicates the attached text traces its derivation from the underlying policy gradient steps.

    Combined views

    7.8K

    1 Source, first seen 29d ago

    Combined views

    7.8K

    1 Source, first seen 29d ago

    182 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    182 likes
    8 comments
    152 saves
    23 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    8 comments
    152 saves
    23 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @cwolferesearchEver wondered where the policy / importance ratio in the PPO loss function comes from? In most explainers of PPO, we first present a vanilla policy gradient (VPG) expression, then go directly to explaining PPO. Usually, PPO is expressed as a loss function that multiplies an importance ratio by the advantage, whereas the VPG takes a product of the gradient of the log probability of an action and some learning signal (e.g., return or advantage). With this in mind, we might wonder: How do we get from this initial VPG expression to what we use in PPO? PPO objective = importance ratio x PG. To understand, we can just take the gradient of the (unclipped) PPO objective. When we do this, we can see that the result is nearly identical to the vanilla policy gradient expression. However, we multiply a standard policy gradient expression by the importance ratio between the current policy and the old policy. Old policy. In PPO, the old policy refers to the policy that is used to sample the rollouts that are being used to compute the policy update in the current batch. Notably, this is different from the reference model, which is used to compute the KL penalty and is usually set equal to the policy before RL training begins. The old policy is different from the current policy because we may perform several sequential policy updates / epochs over sampled data. Importance ratio. Due to these multiple updates, our current policy is actually slightly different from the policy that was used to sample the rollouts. To correct for this mismatch, we can use importance sampling. Formally, importance sampling allows us to estimate an expectation under a target distribution f(x) using samples drawn from a different proposal distribution g(x). Instead of sampling directly from f(x), we can sample from g(x) and correct for the discrepancy between these distributions using the importance ratio f(x) / g(x). This is exactly what we do in PPO, where f(x) is the current policy and g(x) is the old policy. Importance in PPO. In the case of PPO, our importance ratio is the ratio of probabilities for an action between the current and old policy. By multiplying our standard policy gradient expression from the VPG by this ratio, we can correct for the mismatch between current / old policies, allowing us to perform multiple policy updates per batch without disrupting learning.

    1 Source

    @cwolferesearchEver wondered where the policy / importance ratio in the PPO loss function comes from? In most explainers of PPO, we first present a vanilla policy gradient (VPG) expression, then go directly to explaining PPO. Usually, PPO is expressed as a loss function that multiplies an importance ratio by the advantage, whereas the VPG takes a product of the gradient of the log probability of an action and some learning signal (e.g., return or advantage). With this in mind, we might wonder: How do we get from this initial VPG expression to what we use in PPO? PPO objective = importance ratio x PG. To understand, we can just take the gradient of the (unclipped) PPO objective. When we do this, we can see that the result is nearly identical to the vanilla policy gradient expression. However, we multiply a standard policy gradient expression by the importance ratio between the current policy and the old policy. Old policy. In PPO, the old policy refers to the policy that is used to sample the rollouts that are being used to compute the policy update in the current batch. Notably, this is different from the reference model, which is used to compute the KL penalty and is usually set equal to the policy before RL training begins. The old policy is different from the current policy because we may perform several sequential policy updates / epochs over sampled data. Importance ratio. Due to these multiple updates, our current policy is actually slightly different from the policy that was used to sample the rollouts. To correct for this mismatch, we can use importance sampling. Formally, importance sampling allows us to estimate an expectation under a target distribution f(x) using samples drawn from a different proposal distribution g(x). Instead of sampling directly from f(x), we can sample from g(x) and correct for the discrepancy between these distributions using the importance ratio f(x) / g(x). This is exactly what we do in PPO, where f(x) is the current policy and g(x) is the old policy. Importance in PPO. In the case of PPO, our importance ratio is the ratio of probabilities for an action between the current and old policy. By multiplying our standard policy gradient expression from the VPG by this ratio, we can correct for the mismatch between current / old policies, allowing us to perform multiple policy updates per batch without disrupting learning.