Rohan Paul Shares Meta ADeptS-Bench Paper
Tweet flags limits of task success metrics for GUI agents.
TLDR
Rohan Paul posted that a new Meta paper introduces ADeptS-Bench. The benchmark pairs benign and malicious GUI tasks to test seven models. Paul states agents can execute clicks without detecting when the action should never occur. He gives the example of an agent completing a $25K checkout without stopping. The post calls task success alone a dangerously incomplete benchmark for computer-use agents. The attached image shows the first page of the paper.
Combined views
5.4K
2 Sources, first seen 29d ago