FM-Bench Runs Frontier Models as Football Club Managers for 20 Years
Rohan Paul shares details on a benchmark that runs AI agents for two decades in simulation.
TLDR
Rohan Paul posted that current short-term agent benchmarks may miss later failures. He described FM-Bench, which assigns 15 frontier models the task of managing a football club across 20 simulated years and roughly 340 turns. The post states that a model performing well after five years can lag far behind after twenty, and argues that long-running agents require correspondingly long evaluations. The tweet included a screenshot of the arXiv abstract for the benchmark.
Combined views
5.2K
2 Sources, first seen 28d ago