“Coalitional alignment” and reward-model safety/completeness tradeoffs
A research thread describes a condition weaker than requiring reviewer agents to match your utility exactly—and cautions that it is not guaranteed to hold.
TLDR
The thread’s author calls the condition “coalitional alignment,” describing it as substantially weaker than individual alignment, where reviewer agents have exactly your utility. Even that weaker condition is not guaranteed to hold. The author reports some preliminary evidence of non-trivial safety/completeness tradeoffs among real reward models, attributing those effects to coalitional rather than individual alignment.
Combined views
355
1 Source, first seen 15d ago