Safety conditions for one-decision systems and long-running agents
A thread describes a one-decision model where safety holds exactly when the principal's utility is a non-negative weighted combination of reviewer utilities, then extends the characterization to long-running agents.
TLDR
The author describes a precise safety condition for a simple one-decision model: the principal's utility must be expressible as a weighted combination of reviewer utilities, with no negative weights. In that model, the author says, this condition is both necessary and sufficient for safety.
For long-running agents, the thread considers a Markov decision process (MDP), where utilities depend on actions and states, and actions change the state. The author says the performance difference identity extends the one-shot characterization, applied to Q values, to the full MDP.
Combined views
247
1 Source, first seen 16d ago