Roon Argues Simple RLVR Rewards Fail Superintelligent Models
OpenAI technical staff member replies that such functions miss key desiderata and should be banned in favor of model grading.
TLDR
Roon, a pseudonymous technical staff member at OpenAI, replied to users @jachiam0 and @chrislakin that hyper optimization of superintelligent models against simple RLVR reward functions cannot work. He stated the functions contain no terms for the many desiderata people care about. Roon added that these reward functions should be universally banned. He concluded that everything should instead be model graded. The post appears among visible replies on X and reflects his stated position on reward design for advanced AI systems. No further details or responses from the tagged accounts appear in the packet.
Combined views
20.4K
1 Source, first seen 33d ago