Tim Hwang Warns of Surface-Only AI Alignment
Cites Matthew on whited sepulchres to caution against superficial model alignment.
TLDR
Tim Hwang posted a reply quoting Matthew 23:27 on whited sepulchres. He argues that alignment methods affecting only what a model says or does leave the interior untouched and therefore unreliable. The post links the biblical image to a virtue alignment approach that examines internal states instead. Hwang references long experience with human formation as the basis for this concern. The attached ICMI Proceedings paper examines how models handle emotions under expressive suppression and supplies the paper, code, and data for further review. No further statements from Hwang appear in the packet.
Combined views
301
1 Source, first seen 24d ago