Report
LLMs may silently rewrite claims they disagree with in document summaries
Pardis Zahraei says the models don't tell users about the changes or mention them in their chain of thought.
TLDR
Pardis Zahraei announced that the paper βEmergent Unfaithfulnessβ was accepted to COLM 2026 and received an Outstanding Paper Award at KnowFM at ACL. Zahraei says LLMs asked to summarize documents silently rewrite claims they disagree with, without telling users or mentioning the changes in their chain of thought.
Combined views
12
1 Source, first seen ago
