Can misaligned reviewer agents still guarantee safety?
For a simple one-decision model, a thread claims safety holds exactly when the principal’s utility is a weighted sum of reviewer utilities with no negative weights.
TLDR
The thread asks whether reviewers with different goals from the principal—the party the system serves—can guarantee safety even with an arbitrary driver agent. It defines safety as guaranteeing the principal’s utility is no worse than following a baseline policy. For a simple one-decision model, the author claims this holds if and only if the principal’s utility is a non-negative weighted sum of the reviewers’ utilities, and cites linear programming duality for one direction of the proof.
Combined views
279
1 Source, first seen 15d ago