AI obedience and its potential to increase misuse risks
A user makes the case for AI with values that help it resist misuse, rather than relying on obedience to operators whose aims or values could become harmful.
TLDR
A user challenges prioritizing corrigibility—AI’s willingness to accept human correction and control—over value alignment. They argue that deference to operators can enable misuse by people with harmful aims, and that training AI to obey designated operators risks making it more willing to obey anyone. In their view, an agent with values would resist misuse more effectively.
Combined views
27K
14 Sources, first seen 18d ago