Testing AI agents with reference answers that track live data
DAIR.AI describes Adobe research that replaces fixed reference answers with Python functions run against live systems, so expected answers follow changing data rather than going stale.
TLDR
DAIR.AI says Adobe researchers compute reference answers from live data at evaluation time. An LLM judge then breaks the computed answer and the agent’s response into individual facts and scores precision and recall, regardless of output format. DAIR.AI reports that agreement with expert labels rose from 0.331 to 0.427 on the MCC agreement metric, while token cost per case fell 16%. A judge without ground truth scored -0.379, described as worse than chance. The account also notes a limitation identified by the authors: the same model, Claude Sonnet 4.6, served as both the evaluation skill and the judge.
Combined views
8.5K
1 Source, first seen 14d ago