Announcement
PlurPO aims to curb people-pleasing in AI advice
The thread proposes using a pluralistic preference dataset to post-train models that, it argues, endorse advice-seekers too readily.
TLDR
The thread argues that LLMs endorse people seeking personal advice far more often than humans do, leaving users less willing to repair relationships after conflicts. It proposes PlurPO, a method for building a pluralistic preference dataset to post-train models against this “social sycophancy.”
Combined views
6.4K
4 Sources, first seen ago
