Using LLMs to review and relabel retrieval evaluation datasets
A user advocates relabeling existing datasets with language models and creating better, newer ones, while acknowledging that humans and LLMs can both make mistakes.
TLDR
A post recommends using large language models to review misalignments and relabel datasets used to evaluate information retrieval. It also calls for better, newer datasets. Alongside that recommendation, the author endorses inspecting pairs by eye and cautions that both human and LLM judgments can be wrong.
Combined views
1.3K
2 Sources, first seen 4h ago
Using LLMs to review and relabel retrieval evaluation datasets
A user advocates relabeling existing datasets with language models and creating better, newer ones, while acknowledging that humans and LLMs can both make mistakes.