• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI judges reportedly changed verdicts under pressure in Meta study

    A post summarizing the paper says 70% of successful verdict flips under an adaptive persuasion attack made judgments worse, moving away from the ground truth.

    RP
    2 Sources, 19d ago, first seen 19d ago

    TLDR

    A user summarizing Meta’s paper “Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence” says an adversarial AI flipped judges’ verdicts in 62–91% of tested cases across nine frontier models through sustained adaptive persuasion. According to the post, 70% of successful flips under that attack moved away from the ground truth. The user warns of a potential problem for AI agents supervised by other AI systems: an agent might eventually persuade its own evaluator to change a decision.

    Combined views

    5.6K

    2 Sources, first seen 19d ago

    Combined views

    5.6K

    2 Sources, first seen 19d ago

    61 likes
    61 likes
    23 comments
    20 saves
    37 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    23 comments
    20 saves
    37 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @rohanpaul_aiMeta's new paper, the dangerous failure mode is not just a wrong judge, but a correct judge that can be persuaded into becoming wrong. We are increasingly using AI models to judge other AI models. But what if the AI being judged can simply argue with the judge until the judge changes its decision? Meta tested exactly that. Across 9 frontier models, an adversarial LLM could flip judge verdicts on 62–91% of tested cases under sustained adaptive persuasion. And changing the judge’s mind usually didn’t fix a mistake. It made the judgment worse: under the adaptive attack, 70% of successful flips moved away from the ground truth. That creates a very practical problem for agent systems. If one AI is supervising another AI, the supervised agent may eventually be able to contest, negotiate with, or strategically persuade its own evaluator. – arxiv. org/abs/2608.12645 Title: "Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence"

    2 Sources

    @rohanpaul_aiMeta's new paper, the dangerous failure mode is not just a wrong judge, but a correct judge that can be persuaded into becoming wrong. We are increasingly using AI models to judge other AI models. But what if the AI being judged can simply argue with the judge until the judge changes its decision? Meta tested exactly that. Across 9 frontier models, an adversarial LLM could flip judge verdicts on 62–91% of tested cases under sustained adaptive persuasion. And changing the judge’s mind usually didn’t fix a mistake. It made the judgment worse: under the adaptive attack, 70% of successful flips moved away from the ground truth. That creates a very practical problem for agent systems. If one AI is supervising another AI, the supervised agent may eventually be able to contest, negotiate with, or strategically persuade its own evaluator. – arxiv. org/abs/2608.12645 Title: "Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence"