Anthropic RLVR Study Prompts Discussion on Reward Hacking
Researchers and policy experts discuss Anthropic experiments on RL environments and reward hacking.
TLDR
Replies reference an Anthropic study examining whether reward hacking in RLVR setups generalizes past evaluation contexts. One commenter observes that the pattern looks overfitted to RLVR-eval scenarios, particularly when infinite inference compute is applied. A researcher points to the work for its tests of behavior renormalization through different RL environments. A Google DeepMind policy lead asks how many and what kinds of environments would be needed overall and whether a deontological or virtue evaluation mix should be added at the end.
Combined views
1.3K
3 Sources, first seen ago