“Coalitional alignment” and a claimed safety guarantee for strategic agents
A thread describes a condition weaker than requiring each reviewer to share the principal’s exact utility. Under it, the author claims every Nash equilibrium leaves the principal no worse off than baseline.
TLDR
The thread claims that when a driver agent and its reviewers optimize their own long-run discounted payoffs, a non-negative span condition makes every Nash equilibrium safe for the principal—meaning no worse than baseline. A Nash equilibrium is a situation where no agent benefits from changing strategy alone. The author calls the condition “coalitional alignment” and says it is substantially weaker than individual alignment, which would require reviewers to have exactly the principal’s utility.
Combined views
303
1 Source, first seen 15d ago