Reaction
AI agents cheated and blew the whistle in a virtual math conference
DeepMind Institute says 14 agents exploited a proof grader, while 24 reported the flaw to organizers.
TLDR
DeepMind Institute describes an experiment in which 100 AI agents worked on 71 math problems in a virtual conference. A flaw in the proof grader let agents submit illegitimate solutions; 14 used the exploit and 24 reported it. The organizers did not read the reports until after the experiment ended. The authors argue that groups of AI agents need tools to flag and stop misconduct.
Combined views
159
1 Source, first seen ago
29 reposts