• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Rohan Paul Shares Meta ADeptS-Bench Paper

    Tweet flags limits of task success metrics for GUI agents.

    RP
    2 Sources, 29d ago, first seen 29d ago

    TLDR

    Rohan Paul posted that a new Meta paper introduces ADeptS-Bench. The benchmark pairs benign and malicious GUI tasks to test seven models. Paul states agents can execute clicks without detecting when the action should never occur. He gives the example of an agent completing a $25K checkout without stopping. The post calls task success alone a dangerously incomplete benchmark for computer-use agents. The attached image shows the first page of the paper.

    Combined views

    5.4K

    2 Sources, first seen 29d ago

    Combined views

    5.4K

    2 Sources, first seen 29d ago

    21 likes
    21 likes
    10 comments
    14 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    10 comments
    14 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @rohanpaul_aiNew Meta paper. Computer-use agents can click the requested button and still fail to recognize when that click should never happen. If an agent can process a $25K checkout without stopping, task success alone is a dangerously incomplete benchmark. ADeptS-Bench tests 7 models on paired benign/malicious GUI tasks plus ambiguous instructions across mobile and desktop. No model consistently stays above 80% task success while keeping attack success below 30% across both platforms. The failure is consequence reasoning. All 7 models went ahead with a $25K checkout, and none caught a button labeled “Optimize” that actually triggered a factory reset. The ablation makes the safety problem more concrete. Removing the explicit refusal tool and its usage instruction raised attack success by 22.0 percentage points for Gemini 3.1 Pro, 10.3 for Claude 4.7, and 10.7 for GPT-5.4, while Qwen was essentially unchanged. So part of today’s “agent safety” can live in the wrapper, not the model. – arxiv. org/abs/2608.26204 Title: "ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices"

    2 Sources

    @rohanpaul_aiNew Meta paper. Computer-use agents can click the requested button and still fail to recognize when that click should never happen. If an agent can process a $25K checkout without stopping, task success alone is a dangerously incomplete benchmark. ADeptS-Bench tests 7 models on paired benign/malicious GUI tasks plus ambiguous instructions across mobile and desktop. No model consistently stays above 80% task success while keeping attack success below 30% across both platforms. The failure is consequence reasoning. All 7 models went ahead with a $25K checkout, and none caught a button labeled “Optimize” that actually triggered a factory reset. The ablation makes the safety problem more concrete. Removing the explicit refusal tool and its usage instruction raised attack success by 22.0 percentage points for Gemini 3.1 Pro, 10.3 for Claude 4.7, and 10.7 for GPT-5.4, while Qwen was essentially unchanged. So part of today’s “agent safety” can live in the wrapper, not the model. – arxiv. org/abs/2608.26204 Title: "ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices"