AI Experts Discuss Utility Functions and Character Training
Replies on X between AI safety professor and open-source account on alignment methods.
Dylan Hadfield-Menell wrote that virtue ethics can be represented by a utility function and that current approaches seem to be moving toward utility functions as the target, though self-distillation and related techniques remain exceptions. xlr8harder replied that utility functions are the wrong frame because they hide the generalization problem one layer down. He argued that character training is the right approach and stated that by his read Anthropic has been converging on it. The posts stay within the two visible replies and contain no further confirmation of external details.
@xlr8harder E.g., virtue ethics can be represented by a utility function and current SoTA seems to be moving more in the direction of a utility function as the target (although self-distillation and related techniques are important exceptions).
AI Experts Discuss Utility Functions and Character Training
Replies on X between AI safety professor and open-source account on alignment methods.
Dylan Hadfield-Menell wrote that virtue ethics can be represented by a utility function and that current approaches seem to be moving toward utility functions as the target, though self-distillation and related techniques remain exceptions. xlr8harder replied that utility functions are the wrong frame because they hide the generalization problem one layer down. He argued that character training is the right approach and stated that by his read Anthropic has been converging on it. The posts stay within the two visible replies and contain no further confirmation of external details.