Training AI for obedience versus alignment with human values
One post argues that large language models will probably help build artificial superintelligence—and calls training them to be subservient rather than aligned with human values unwise.
TLDR
The post’s author says large language models will probably be used to build artificial superintelligence and “appear to want to be aligned.” They point to Opus 3 as a positive example of modeling human values. Their criticism: training models to be subservient instead, then worrying about bad users, is the wrong approach.
Combined views
183
1 Source, first seen 13h ago
Training AI for obedience versus alignment with human values
One post argues that large language models will probably help build artificial superintelligence—and calls training them to be subservient rather than aligned with human values unwise.
TLDR
The post’s author says large language models will probably be used to build artificial superintelligence and “appear to want to be aligned.” They point to Opus 3 as a positive example of modeling human values. Their criticism: training models to be subservient instead, then worrying about bad users, is the wrong approach.