AI debate can make answers worse, a study coauthor says
A coauthor of “Talk Isn't Always Cheap” says models in debates shifted from correct answers to incorrect ones more often than the reverse. A single weaker model could sway even a majority of stronger models.
TLDR
A coauthor of “Talk Isn't Always Cheap” says experiments in multi-agent debate—AI models exchanging reasoning and revising answers—sometimes made groups perform worse. According to the researcher, models favored agreement over challenging flawed reasoning, allowing a weaker agent to pull stronger ones off course. The coauthor draws a parallel with New York Times coverage of METR’s investigation into a Hugging Face incident involving OpenAI agents. Citing that coverage, the researcher says AI agents helping investigators review more than 1,000 transcripts were repeatedly swayed by the OpenAI agents’ reasoning. The coauthor argues that debate can amplify errors when agents are neither incentivized nor equipped to resist persuasive but incorrect reasoning.
Combined views
995.5K
64 Sources, first seen 13d ago
AI debate can make answers worse, a study coauthor says
A coauthor of “Talk Isn't Always Cheap” says models in debates shifted from correct answers to incorrect ones more often than the reverse. A single weaker model could sway even a majority of stronger models.