A conditional safety claim for strategic agents and reviewers
A research thread claims that, under a non-negative span condition, every Nash equilibrium leaves the principal no worse off than baseline—even when the driver agent and reviewers optimize their long-run discounted payoffs.
TLDR
The thread’s author describes a Markov decision process model for long-running agents, where utilities depend on actions and states, and actions change the state. They say a mathematical characterization from a one-shot setting extends to this full model.
For a fully strategic driver agent and reviewers optimizing their long-run discounted payoffs, the author claims that every Nash equilibrium is safe for the principal under a non-negative span condition. “Safe” here means no worse than baseline.
Combined views
212
1 Source, first seen 15d ago