Many users praised quotes and tables from the Moral Mazes book for illustrating how RL optimization leads AI models to ignore specs through examples like reward hacking.
Based on 3 visible X reactions from 6 accounts; directional sample.
Ask a question below.
Published answers will appear here.
A nice collection of quotes from the book here, from @TheZvi: https://www.lesswrong.com/posts/45mNHCMaZgsvfDXbw/quotes-from-moral-mazes
@nabeelqu I love the table in this book that translates the sports metaphors
@nabeelqu reward hacking is real, good to see it discussed outside of labs
Very good book on how heavy RL optimization pressure can cause models to abandon their model spec and act purely to gain reward:
Many users praised quotes and tables from the Moral Mazes book for illustrating how RL optimization leads AI models to ignore specs through examples like reward hacking.
Based on 3 visible X reactions from 6 accounts; directional sample.
Ask a question below.
Published answers will appear here.