Report
Leaderboard scores and the risk of choosing a worse model
Posts describing new research argue that leaderboards can report progress while steering people toward a worse model. They point to LiveBench returning 18 task scores instead of one as an example of score granularity.
TLDR
Posts describing new research argue that leaderboard results can show progress while steering model selection toward a worse model. They highlight LiveBench’s 18 task scores instead of one and say selection-loss and generalization-gap curves come much closer together when released scores are counted rather than submissions.
Combined views
281
10 Sources, first seen 8h ago
