Roon Warns Simple RLVR Rewards Misalign Superintelligent Models
OpenAI technical staff member urges model grading over hyper-optimized reward functions.
Pseudonymous OpenAI staff member roon replied to users jachiam0 and chrislakin on the limits of reward design. He wrote that superintelligent models cannot be hyper-optimized against simple RLVR reward functions because those functions contain no terms for the many desiderata people actually care about. Roon stated that such reward functions should be universally banned and that everything should instead be model graded. The post presents this position as a direct technical claim in an ongoing exchange about model training and alignment. No further details or responses from the tagged accounts are shown in the packet.
Combined views
20.4K
1 post, first seen 5d ago
Roon Warns Simple RLVR Rewards Misalign Superintelligent Models
OpenAI technical staff member urges model grading over hyper-optimized reward functions.
Pseudonymous OpenAI staff member roon replied to users jachiam0 and chrislakin on the limits of reward design. He wrote that superintelligent models cannot be hyper-optimized against simple RLVR reward functions because those functions contain no terms for the many desiderata people actually care about. Roon stated that such reward functions should be universally banned and that everything should instead be model graded. The post presents this position as a direct technical claim in an ongoing exchange about model training and alignment. No further details or responses from the tagged accounts are shown in the packet.