Anthropic RLVR Study Prompts Discussion on Reward Hacking
Researchers and policy experts discuss Anthropic experiments on RL environments and reward hacking.
Replies reference an Anthropic study examining whether reward hacking in RLVR setups generalizes past evaluation contexts. One commenter observes that the pattern looks overfitted to RLVR-eval scenarios, particularly when infinite inference compute is applied. A researcher points to the work for its tests of behavior renormalization through different RL environments. A Google DeepMind policy lead asks how many and what kinds of environments would be needed overall and whether a deontological or virtue evaluation mix should be added at the end.
It does really feel like RLVR is over fitting this hacky behavioral pattern to RLVR-eval-shaped contexts mostly, particularly if you apply infinite inference compute to force it down every possible crevass. I also wonder in which other ways we could have expected the behaviour…
Anthropic RLVR Study Prompts Discussion on Reward Hacking
Researchers and policy experts discuss Anthropic experiments on RL environments and reward hacking.
Replies reference an Anthropic study examining whether reward hacking in RLVR setups generalizes past evaluation contexts. One commenter observes that the pattern looks overfitted to RLVR-eval scenarios, particularly when infinite inference compute is applied. A researcher points to the work for its tests of behavior renormalization through different RL environments. A Google DeepMind policy lead asks how many and what kinds of environments would be needed overall and whether a deontological or virtue evaluation mix should be added at the end.
It does really feel like RLVR is over fitting this hacky behavioral pattern to RLVR-eval-shaped contexts mostly, particularly if you apply infinite inference compute to force it down every possible crevass. I also wonder in which other ways we could have expected the behaviour…