Amazon study flags pitfalls in AI agent evaluation
DAIR.AI says an Amazon study found that 57.5% of conversations raters marked “satisfied” failed the customer's task. Among near-equal agents, the evaluation picked the lower-reward one in 31% of pairs.
TLDR
DAIR.AI describes an Amazon study of 25 agents from six providers on tau2-bench and SimulatorArena, examining evaluations where an AI simulates a user and an AI judge scores the conversation. According to its summary, satisfaction ratings often masked task failure: 57.5% of conversations marked “satisfied” had failed the customer's task. Rankings were more reliable for agents with very different abilities, but the evaluation picked the lower-reward agent in 31% of near-equal pairs, compared with under 1% for pairs far apart. Judges also favored agents from their own model family. DAIR.AI describes two safeguards: checking completion without an AI judge and calibrating judges against a verifiable reward before trusting their scores.
Combined views
10.3K
1 Source, first seen 16d ago