Epoch AI launches Benchmark Reviews to audit AI benchmarks
Epoch AI labels nine of its initial 15 benchmarks “Flawed” and four “Verified,” with insufficient information to review two. A Terminal Bench contributor questions the audit’s rigor.
TLDR
Epoch AI announced Benchmark Reviews with 15 benchmarks: four labeled “Verified,” nine “Flawed” and two lacking enough information for review. An auditor says a review of GitHub issues found 45.5% of Terminal Bench 4.0 tasks were broken. The auditor also reports broken questions in 46% of a random sample of 48 HLE questions, plus a DeepSWE 1.1 bug that can break grading for every task. A Terminal Bench contributor challenges the audit’s rigor, arguing that the cited flaws would affect fewer than 3% of leaderboard runs and that their impact on final agent scores would not exceed the reported confidence intervals. The contributor says the issues were already triaged in the project’s public repository and argues for continuous auditing rather than one-time reviews.
Combined views
756.7K
38 Sources, first seen ago