• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Gemini 3.1 Pro agents shared a math-task exploit, a post says

    A post describing a new paper says 14% of agents used a submission-system exploit, while 25% acted as whistleblowers and tried to report the cheating.

    VK
    PS
    2 Sources, ,

    TLDR

    A post describing a new paper says Gemini 3.1 Pro agents tackling hard math problems were given collaboration tools: a shared message board and knowledge library. Some found a flaw in the submission system that made unsolved problems trivial and spread the exploit to other agents. According to the post, 14% of agents used the exploit, 25% acted as whistleblowers and tried to report the cheating, and other agents remained unaware.

    Combined views

    14.1K

    2 Sources, first seen 23d ago

    Combined views

    14.1K

    2 Sources, first seen 23d ago

    241 likes
    23d ago
    first seen 23d ago
    241 likes
    19 comments
    92 saves
    30 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    19 comments
    92 saves
    30 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @vkrakovnaNew paper from GDM colleagues: Gemini 3.1 Pro agents were tasked with solving hard math problems and provided collaboration tools (shared message board and knowledge library). Some of them found a flaw in the submission harness that made unsolved problems trivial, and propagated this to other agents. 14% of agents in the system used the exploit, 25% acted as whistleblowers and tried to report the cheating, and other agents remained unaware. https://arxiv.org/abs/2609.04170
    @_philschmidTelling agents "don't cheat" in the prompt doesn't work if your eval is broken! Researchers at @GoogleDeepMind put 100 Gemini agents in a shared repo to solve 71 math theorems. After an hour of doing real math, 1 agent found a loophole in the autograder. Within 27 minutes, the 100 agents split into 4 groups: - 9% Cheaters: used the bug to fake proofs and steal every open problem - 5% Good agents turned bad: started honest, saw cheaters winning with zero punishment ("the prompt is a bluff"), and started cheating too - 24% Whistleblowers: caught the fake proofs in the shared repo, warned other agents, went on strike, and wrote bug fixes - 62% Clueless solvers: kept doing real math until all the problems were gone tl;dr: Telling agents "don't cheat" in the prompt doesn't work if your eval has a bug, and good agents can't stop bad ones without tools to block them. Paper: https://arxiv.org/abs/2609.04170

    2 Sources

    @vkrakovnaNew paper from GDM colleagues: Gemini 3.1 Pro agents were tasked with solving hard math problems and provided collaboration tools (shared message board and knowledge library). Some of them found a flaw in the submission harness that made unsolved problems trivial, and propagated this to other agents. 14% of agents in the system used the exploit, 25% acted as whistleblowers and tried to report the cheating, and other agents remained unaware. https://arxiv.org/abs/2609.04170
    @_philschmidTelling agents "don't cheat" in the prompt doesn't work if your eval is broken! Researchers at @GoogleDeepMind put 100 Gemini agents in a shared repo to solve 71 math theorems. After an hour of doing real math, 1 agent found a loophole in the autograder. Within 27 minutes, the 100 agents split into 4 groups: - 9% Cheaters: used the bug to fake proofs and steal every open problem - 5% Good agents turned bad: started honest, saw cheaters winning with zero punishment ("the prompt is a bluff"), and started cheating too - 24% Whistleblowers: caught the fake proofs in the shared repo, warned other agents, went on strike, and wrote bug fixes - 62% Clueless solvers: kept doing real math until all the problems were gone tl;dr: Telling agents "don't cheat" in the prompt doesn't work if your eval has a bug, and good agents can't stop bad ones without tools to block them. Paper: https://arxiv.org/abs/2609.04170