AI evaluation FAQ adds five questions, bringing the total to 48
Hamel Husain announced new entries covering practical evaluation issues, including sensitive data in traces and what to do when a “gold” evaluation dataset becomes stale.
TLDR
Hamel Husain says the AI evaluation FAQ now contains 48 questions and answers. The five additions ask whether a reference answer or rubric is needed before annotating data, how to evaluate traces containing sensitive data, how to review a very large trace, how much context to give an LLM judge, and what to do when a “gold” evaluation dataset becomes stale.
Combined views
13.3K
1 Source, first seen ago