• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Testing AI agents with reference answers that track live data

    DAIR.AI describes Adobe research that replaces fixed reference answers with Python functions run against live systems, so expected answers follow changing data rather than going stale.

    DA
    1 Source, 14d ago, first seen 14d ago

    TLDR

    DAIR.AI says Adobe researchers compute reference answers from live data at evaluation time. An LLM judge then breaks the computed answer and the agent’s response into individual facts and scores precision and recall, regardless of output format. DAIR.AI reports that agreement with expert labels rose from 0.331 to 0.427 on the MCC agreement metric, while token cost per case fell 16%. A judge without ground truth scored -0.379, described as worse than chance. The account also notes a limitation identified by the authors: the same model, Claude Sonnet 4.6, served as both the evaluation skill and the judge.

    Combined views

    8.5K

    1 Source, first seen 14d ago

    Combined views

    8.5K

    1 Source, first seen 14d ago

    76 likes
    76 likes
    14 comments
    74 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    14 comments
    74 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @dair_aiNice work from Adobe. It's standard practice to store a fixed reference answer for every eval case. When the underlying data changes daily, that stored answer goes stale. Adobe researchers write each reference answer as a Python function instead. The function runs against the live system at evaluation time, so the expected answer follows the data, and an upstream API change makes the test fail visibly. An LLM judge then splits the agent's response and the computed answer into atomic facts and scores precision and recall, whatever the output format. Against expert labels, this raises agreement from an MCC of 0.331 to 0.427 and cuts token cost per case by 16%. A judge working with no ground truth scored an MCC of -0.379, which is worse than chance. The pipeline runs as a harness skill. The authors list one limitation, which is that the same model, Claude Sonnet 4.6, acted as both the skill and the judge. Paper: https://academy.dair.ai/papers/skill-based-agentic-evaluation-for-real-time-data-science-tasks-2609.16487

    1 Source

    @dair_aiNice work from Adobe. It's standard practice to store a fixed reference answer for every eval case. When the underlying data changes daily, that stored answer goes stale. Adobe researchers write each reference answer as a Python function instead. The function runs against the live system at evaluation time, so the expected answer follows the data, and an upstream API change makes the test fail visibly. An LLM judge then splits the agent's response and the computed answer into atomic facts and scores precision and recall, whatever the output format. Against expert labels, this raises agreement from an MCC of 0.331 to 0.427 and cuts token cost per case by 16%. A judge working with no ground truth scored an MCC of -0.379, which is worse than chance. The pipeline runs as a harness skill. The authors list one limitation, which is that the same model, Claude Sonnet 4.6, acted as both the skill and the judge. Paper: https://academy.dair.ai/papers/skill-based-agentic-evaluation-for-real-time-data-science-tasks-2609.16487