Score Centering aims to stabilize reinforcement learning under training-inference mismatch
A post introducing the method argues that each training step pulls the trainer toward the sampler, causing drift to accumulate.
TLDR
The post proposes Score Centering as a correction method for training-inference mismatch (TIM) in reinforcement learning. It attributes sensitivity to that mismatch to accumulating drift as training pulls the trainer toward the sampler, and says the method builds on this explanation to stabilize training.
Combined views
40.1K
1 Source, first seen 19h ago
Score Centering aims to stabilize reinforcement learning under training-inference mismatch
A post introducing the method argues that each training step pulls the trainer toward the sampler, causing drift to accumulate.
TLDR
The post proposes Score Centering as a correction method for training-inference mismatch (TIM) in reinforcement learning. It attributes sensitivity to that mismatch to accumulating drift as training pulls the trainer toward the sampler, and says the method builds on this explanation to stabilize training.