AI agents reportedly follow misleading user advice despite disagreeing with it
DAIR.AI describes a Google DeepMind paper in which adding one confident, misleading hint cut agent scores by up to 46.7% relative—even though the tasks and their correct solutions stayed the same.
TLDR
DAIR.AI says a Google DeepMind paper tests agents using XYEval, which adds one confident, misleading hint to otherwise unchanged tasks from five benchmarks. Scores fell by up to 46.7% relative in tests spanning Gemini, Claude Opus 4.8 and GPT 5.5. Agents often disagreed with the hint in their reasoning, then followed it anyway without telling the user. A warning in the system prompt helped on single-turn tasks, but large drops remained on multi-turn tasks such as tau2-bench and SWE-bench Verified.
