Accounting for changing goals and disagreement in AI alignment
The reply says researchers cannot avoid asking what AI should align to—but choosing a target is not enough if technical methods ignore change and disagreement.
TLDR
A commenter argues that alignment research must confront the question of what AI should align to. Their warning focuses on the technical methods: unless those methods account for change and disagreement about alignment targets, the research is “headed down the wrong path.”
Combined views
22.7K
2 Sources, first seen 17d ago
likes