Researcher Questions Post-Training Benchmark Reliability
Xiuyu Li notes how evaluations can be manipulated through settings and harness choices.
TLDR
Xiuyu Li, a researcher at StepFun_ai, posted reflections on a Zhihu discussion about post-training benchmarks. He stated that benchmarks saturate over time and that reasoning and coding evaluations are growing more generalized and crowdsourced. Li described LLM evaluations as makeshift arrangements where companies can choose their own settings, including temperature, top_p, length, or harness parameters, leaving room for adjustments. Songlin Yang retweeted the post without added comment.
Combined views
39.1K
2 Sources, first seen 26d ago