AI evaluation FAQ adds five questions, bringing the total to 48
Hamel Husain announced new entries covering practical evaluation issues, including sensitive data in traces and what to do when a “gold” evaluation dataset becomes stale.
TLDR
Hamel Husain says the AI evaluation FAQ now contains 48 questions and answers. The five additions ask whether a reference answer or rubric is needed before annotating data, how to evaluate traces containing sensitive data, how to review a very large trace, how much context to give an LLM judge, and what to do when a “gold” evaluation dataset becomes stale.
Combined views
13.3K
1 Source, first seen 2d ago
AI evaluation FAQ adds five questions, bringing the total to 48
Hamel Husain announced new entries covering practical evaluation issues, including sensitive data in traces and what to do when a “gold” evaluation dataset becomes stale.
TLDR
Hamel Husain says the AI evaluation FAQ now contains 48 questions and answers. The five additions ask whether a reference answer or rubric is needed before annotating data, how to evaluate traces containing sensitive data, how to review a very large trace, how much context to give an LLM judge, and what to do when a “gold” evaluation dataset becomes stale.