The possibility of ‘spiky’ AI alignment across domains
One post predicts models will probably be very well aligned in some domains and extremely misaligned in others simultaneously. A quote-post calls that a fact.
TLDR
One post argues that AI alignment is likely to be “spiky,” much like model capabilities: a model could be very well aligned in some domains while extremely misaligned in others. The person quoting it rejects the tentative framing, saying, “no, this is a fact.”
The possibility of ‘spiky’ AI alignment across domains
One post predicts models will probably be very well aligned in some domains and extremely misaligned in others simultaneously. A quote-post calls that a fact.
TLDR
One post argues that AI alignment is likely to be “spiky,” much like model capabilities: a model could be very well aligned in some domains while extremely misaligned in others. The person quoting it rejects the tentative framing, saying, “no, this is a fact.”