Can an agent system guarantee safety when reviewers’ goals differ from the user’s?
A user studying reviewer agents defines safety as guaranteeing that the user’s utility—what they want to maximize—is no worse than under a baseline policy.
TLDR
A user describes a model where a driver agent repeatedly proposes actions. Reviewer agents compare those proposals with a baseline and vote to approve or deny them based on perceived utility. The challenge is that reviewers’ utility functions differ from the user’s. The author asks whether the system can still guarantee that the user does no worse than following the baseline policy, even with an arbitrary driver agent.
Combined views
528
1 Source, first seen 16d ago