• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    FM-Bench Runs Frontier Models as Football Club Managers for 20 Years

    Rohan Paul shares details on a benchmark that runs AI agents for two decades in simulation.

    RP
    2 Sources, 28d ago, first seen 28d ago

    TLDR

    Rohan Paul posted that current short-term agent benchmarks may miss later failures. He described FM-Bench, which assigns 15 frontier models the task of managing a football club across 20 simulated years and roughly 340 turns. The post states that a model performing well after five years can lag far behind after twenty, and argues that long-running agents require correspondingly long evaluations. The tweet included a screenshot of the arXiv abstract for the benchmark.

    Combined views

    5.2K

    2 Sources, first seen 28d ago

    Combined views

    5.2K

    2 Sources, first seen 28d ago

    18 likes
    18 likes
    9 comments
    5 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    9 comments
    5 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @rohanpaul_aiCurrent agent benchmarks may be ending before the real failures start. FM-Bench shows that a model can look strong after 5 years and still finish far behind after 20, so long-running agents need long-running evaluations. FM-Bench has 15 frontier models manage a football club for 20 simulated years, across roughly 340 to 400 decision stops where transfers, contracts, cash, investments, and rival actions keep changing the future. The rankings barely resemble their final shape early on. On seed 1, the year-5 ranking correlated just 0.19 with the final order, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th. Competition changes the picture too. In the shared Arena, 10 different models won the league at least once instead of one early leader simply compounding forever. What tracked stronger performance was managerial behavior: cutting slow-payoff investments near the end, keeping cash deployed, and renewing contracts earlier. Token use spanned about 7X and still did not order the board. So for long-running agents, short task success is a weak proxy for sustained decision quality. The caveat is that the solo board uses 3 seeds and the Arena only 1 shared world. – arxiv. org/abs/2608.18423 Title: "FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents"

    2 Sources

    @rohanpaul_aiCurrent agent benchmarks may be ending before the real failures start. FM-Bench shows that a model can look strong after 5 years and still finish far behind after 20, so long-running agents need long-running evaluations. FM-Bench has 15 frontier models manage a football club for 20 simulated years, across roughly 340 to 400 decision stops where transfers, contracts, cash, investments, and rival actions keep changing the future. The rankings barely resemble their final shape early on. On seed 1, the year-5 ranking correlated just 0.19 with the final order, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th. Competition changes the picture too. In the shared Arena, 10 different models won the league at least once instead of one early leader simply compounding forever. What tracked stronger performance was managerial behavior: cutting slow-payoff investments near the end, keeping cash deployed, and renewing contracts earlier. Token use spanned about 7X and still did not order the board. So for long-running agents, short task success is a weak proxy for sustained decision quality. The caveat is that the solo board uses 3 seeds and the Arena only 1 shared world. – arxiv. org/abs/2608.18423 Title: "FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents"