RL Training Creates Split Personas In LLMs Across Environments
Seth Lazar notes supporting evidence from a paper under revision on incoherent values.
A retweet from Dylan Hadfield-Menell flags a new LessWrong post. That post states reinforcement learning causes large language models to develop split personas whose propensities, values, and beliefs differ across environments. Seth Lazar replies that his team's incoherent values paper contains some evidence that potentially confirms the claim. Lazar adds that the paper is now being revised. The visible posts limit discussion to these statements from the LessWrong argument and the direct reply referencing the ongoing research.
Combined views
6.1K
3 posts, first seen 4d ago