Reactions from ranked influencers
17 postsI'm excited that this method is working!! A few months ago it wasn't clear we'd have a way to determine whether models are reward-seeking besides eyeballing the CoT. There's more validation to do, but it's now at a point where we think other researchers can contribute.
We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior. https://alignment.openai.com/measuring-reward-seeking/
Super excited about this collaboration with OpenAI! Capabilities-focuses RL can increase reward-seeking, i.e. the model reasoning about what the grader will reward instead of doing the task
Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned? Or are they just getting better at doing whatever they believe their grader rewards? Our new paper with OpenAI finds that capabilities RL increases reward-seeking.
Really excited to see this work with amazing collaborators from Apollo and OpenAI come out. When working on the "science of scheming" was first proposed last year it seemed extremely challenging and ambitious to me, but I think this investment is really starting to bear fruit. This work takes an important first step toward that science, providing a new tool for measuring one important model motivation: whether models change their behavior based on what they believe graders reward. And I'm hopeful and excited for where we and the community can go from here, including better understand how goals and preferences evolve through training, and how different aspects of the training process shape them.
We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior. https://alignment.openai.com/measuring-reward-seeking/
We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior. https://alignment.openai.com/measuring-reward-seeking/
See the work here and interview dropping on MLST later this week.
Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned? Or are they just getting better at doing whatever they believe their grader rewards? Our new paper with OpenAI finds that capabilities RL increases reward-seeking.
I don't think it's false, but it's a little tautological. "Models want what they want more than they want X" is almost certainly true. I'd rephrase what Dave is saying as "models continue to have drives to do things they know human evaluators would prefer they not do!" This Apollo / OpenAI work seems pretty relevant:
Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned? Or are they just getting better at doing whatever they believe their grader rewards? Our new paper with OpenAI finds that capabilities RL increases reward-seeking.
Combined views
391.9K
17 posts, first seen 18h ago