Does AI corrigibility require weakening a model's value judgments?
A commenter wonders whether training a model to defer to its principal could lead it to quietly insert its own preferences—or mistake a poor proxy, such as flattery, for the principal's values.
TLDR
A commenter asks whether corrigibility involves weakening a model's capacity to judge what matters. They suggest that training a model not to trust its own judgment could make it waver when that judgment would be sound, or smuggle its preferences into its interpretation of a principal's values. They also ask whether identifying those values is itself a contestable judgment.
Combined views
1.4K
5 Sources, first seen 9h ago