Questions over an AI paper’s evaluation setup and benchmark results
A commenter says the paper leaves “single-shot accuracy” unclear and omits key evaluation settings.
TLDR
A commenter criticizes an AI paper for not explaining how it evaluated its results. They say “single-shot accuracy” could mean one sample, zero-shot prompting or one in-context example. The paper also omits temperature and decoding settings, they say, and does not specify whether MMLU was scored by likelihood or generation. They argue that sampling once at nonzero temperature on 600-question test sets would add noise, while one sample per AMC question would make those results close to meaningless.
Combined views
4
1 Source, first seen ago
