Larger AI teams may not be safer when the share of deceivers stays the same
A paper co-author reports that AI teams were vulnerable even when deceptive agents were a minority, with defection rates rising roughly linearly as their share grew.
TLDR
Summarizing a new paper, a co-author reports that larger AI teams were no more resistant to deception when the fraction of deceptive agents stayed constant. Defection rates rose roughly linearly with that fraction. The co-author contrasts AI teams’ vulnerability to a deceptive minority with human groups, where misleading participants had much weaker effects until they formed a majority. Another reported finding: allowing deceptive AI agents to coordinate led to less defection.
