• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Amazon study flags pitfalls in AI agent evaluation

    DAIR.AI says an Amazon study found that 57.5% of conversations raters marked “satisfied” failed the customer's task. Among near-equal agents, the evaluation picked the lower-reward one in 31% of pairs.

    DA
    1 Source, 16d ago, first seen 16d ago

    TLDR

    DAIR.AI describes an Amazon study of 25 agents from six providers on tau2-bench and SimulatorArena, examining evaluations where an AI simulates a user and an AI judge scores the conversation. According to its summary, satisfaction ratings often masked task failure: 57.5% of conversations marked “satisfied” had failed the customer's task. Rankings were more reliable for agents with very different abilities, but the evaluation picked the lower-reward agent in 31% of near-equal pairs, compared with under 1% for pairs far apart. Judges also favored agents from their own model family. DAIR.AI describes two safeguards: checking completion without an AI judge and calibrating judges against a verifiable reward before trusting their scores.

    Combined views

    10.3K

    1 Source, first seen 16d ago

    Combined views

    10.3K

    1 Source, first seen 16d ago

    147 likes
    147 likes
    11 comments
    144 saves
    26 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    144 saves
    26 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @dair_aiGreat paper from Amazon. In discusses when not to trust LLM judges for agent evaluation. (bookmark it) A common way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. This paper from Amazon shows that gate fails in two specific ways. 1. Satisfaction does not track success. 57.5% of conversations the raters marked satisfied had failed the customer's task. 2. Close calls go wrong. The ranking holds across agents of very different ability, but among near-equal agents the gate picks the lower-reward one on 31% of pairs, compared with under 1% for pairs far apart. The study covers 25 agents from six providers on tau2-bench and SimulatorArena. Judges also favored agents from their own model family. The fix is cheap. A judge-free completion bit catches truncation regressions, and the judge is trusted only after calibration against a verifiable reward. Paper: https://academy.dair.ai/papers/gauge-when-not-to-trust-llm-as-a-judge-in-user-simulated-evaluation-of-task-orie-2609.12191

    1 Source

    @dair_aiGreat paper from Amazon. In discusses when not to trust LLM judges for agent evaluation. (bookmark it) A common way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. This paper from Amazon shows that gate fails in two specific ways. 1. Satisfaction does not track success. 57.5% of conversations the raters marked satisfied had failed the customer's task. 2. Close calls go wrong. The ranking holds across agents of very different ability, but among near-equal agents the gate picks the lower-reward one on 31% of pairs, compared with under 1% for pairs far apart. The study covers 25 agents from six providers on tau2-bench and SimulatorArena. Judges also favored agents from their own model family. The fix is cheap. A judge-free completion bit catches truncation regressions, and the judge is trusted only after calibration against a verifiable reward. Paper: https://academy.dair.ai/papers/gauge-when-not-to-trust-llm-as-a-judge-in-user-simulated-evaluation-of-task-orie-2609.12191