A proposed AI reward signal for a persistence–trustworthiness tradeoff
A user argues that a tradeoff they say Saachi Jain and OpenAI reported to the press stems from a scalar, totally ordered reward signal. They believe a preference-based alternative could avoid it at a small additional compute cost.
TLDR
The user proposes generating concurrent, independent model rollouts, then having the same model checkpoint judge them using detailed rubrics and map its overall preferences in a Hasse diagram. They believe this reward signal could avoid a persistence–trustworthiness tradeoff they attribute to scalar rewards, at a small additional compute cost. They say Saachi Jain and OpenAI reported the tradeoff to the press.
Combined views
824
1 Source, first seen 2h ago
