Reported problems in the MMLU AI benchmark
A contributor to MMLU-Redux says manual annotation of many MMLU subsets uncovered numerous issues, especially in virology.
TLDR
An MMLU-Redux contributor says the team manually annotated many MMLU subsets and found numerous issues, particularly in virology. The linked paper, “Are We Done with MMLU?”, examines errors in the widely used Massive Multitask Language Understanding benchmark. The contributor recommended it in a reply to a post praising Epoch’s public AI benchmarking while arguing that its work highlights the poor quality of some popular benchmarks.
Combined views
96
1 Source, first seen 22h ago
Reported problems in the MMLU AI benchmark
A contributor to MMLU-Redux says manual annotation of many MMLU subsets uncovered numerous issues, especially in virology.
TLDR
An MMLU-Redux contributor says the team manually annotated many MMLU subsets and found numerous issues, particularly in virology. The linked paper, “Are We Done with MMLU?”, examines errors in the widely used Massive Multitask Language Understanding benchmark. The contributor recommended it in a reply to a post praising Epoch’s public AI benchmarking while arguing that its work highlights the poor quality of some popular benchmarks.