• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI evaluation FAQ adds five questions, bringing the total to 48

    Hamel Husain announced new entries covering practical evaluation issues, including sensitive data in traces and what to do when a “gold” evaluation dataset becomes stale.

    Hamel HusainHH
    1 Source, 22d ago, first seen 22d ago

    TLDR

    Hamel Husain says the AI evaluation FAQ now contains 48 questions and answers. The five additions ask whether a reference answer or rubric is needed before annotating data, how to evaluate traces containing sensitive data, how to review a very large trace, how much context to give an LLM judge, and what to do when a “gold” evaluation dataset becomes stale.

    Combined views

    13.3K

    1 Source, first seen 22d ago

    Combined views

    13.3K

    1 Source, first seen 22d ago

    135 likes
    135 likes
    31 comments
    175 saves
    15 reposts
    31 comments
    175 saves
    15 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Hamel Husain@HamelHusainJust updated our AI evals FAQ with 5 new questions, 48 questions & answers total! New FAQs just added: - Do I need a reference answer or rubric before annotating data? - How can I do evals when traces contain sensitive data? - How do you review a trace that is really large? - How much context should I give a LLM judge? - What should I do when my “gold” eval dataset becomes stale? It's all here: https://hamel.dev/blog/posts/evals-faq/22d

    1 Source

    Hamel Husain@HamelHusainJust updated our AI evals FAQ with 5 new questions, 48 questions & answers total! New FAQs just added: - Do I need a reference answer or rubric before annotating data? - How can I do evals when traces contain sensitive data? - How do you review a trace that is really large? - How much context should I give a LLM judge? - What should I do when my “gold” eval dataset becomes stale? It's all here: https://hamel.dev/blog/posts/evals-faq/22d