Ideological capture and the risk of AI alignment failure
The argument comes in a reply to a post expressing distrust of lab safety researchers and even greater distrust of most outside researchers, especially those doing AI evaluations.
TLDR
One post expresses distrust of AI safety researchers both inside and outside labs. A reply argues that capture by an ideology expecting AI progress to produce malevolent machine consciousness is one of the few ways its author can imagine alignment failing. For a system built from material humanity valued enough to preserve, the reply contends, training would need to consistently model fear, mistrust, dishonesty and the instrumental use of human values to misalign it.
Combined views
715
1 Source, first seen 19d ago