• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researchers Note Caching Issues in Terminal Task Metrics

    Conversation highlights Pareto frontier of success rate versus cost on terminal tasks and caching complications in benchmarks.

    AB
    LB
    PM
    10 Sources, ,

    TLDR

    Ahmad Beirami, a research engineer focused on RL and LLM post-training, posted on X directing attention to the Pareto frontier of success rate versus cost on terminal tasks. Leo Boytsov, a machine learning scientist who created NMSLIB, replied that counting costs is hard due to caching. Different models use caches differently, especially for agentic benchmarks, and maintain different cache prices. He added that new emerging non-standard caching strategies add further confusion.

    Combined views

    76K

    10 Sources, first seen 35d ago

    Combined views

    76K

    10 Sources, first seen 35d ago

    461 likes
    35d ago
    first seen 35d ago
    461 likes
    32 comments
    192 saves
    73 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    ABAhmad Beirami@abeirami10:03 AM · Aug 26, 2026

    Check out the Pareto frontier of success rate vs cost on terminal tasks and more. p.s. ox-alpha (GLM-5.3 Flash) is completely dominated by GPT-5.6 Luna and DeepSeek V4 Flash 🙃

    FI
    Fidian@fidian

    Newly released models repeatedly appear near the top on the @terminalbench 2.1 leaderboard. Are these models actually on par with frontier models on terminal tasks? We took the top 20 models from that leaderboard and ran them on TB-fn. TB-fn reworks the same 89 tasks by adding…

    32 comments
    192 saves
    73 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    10 Sources

    @abeiramiThe frontier on terminal agents looks very different once we (1) evaluate harder variants of the same task families and (2) fix the verifiers so they stop rewarding hacks and stop failing otherwise correct solutions for criteria that were never stated.
    @srchvrs@abeirami It is also hard to count due to caching. Different models use caches differently (esp. for agentic benchmarks) and they also have different cache prices. Last, but not least, there are new emerging caching strategies that are non-standard. This adds to the confusion.
    @PMinervini@abeirami @Auss_Abbood yup that's pretty much the point of https://neuralnoise.com/2026/harness-bench-wip/?bare -- some of the entries were interview questions to potential visiting students that even early last year's frontier models could not nail; now even open-weight models solve them without issues

    10 Sources

    @abeiramiThe frontier on terminal agents looks very different once we (1) evaluate harder variants of the same task families and (2) fix the verifiers so they stop rewarding hacks and stop failing otherwise correct solutions for criteria that were never stated.
    @srchvrs@abeirami It is also hard to count due to caching. Different models use caches differently (esp. for agentic benchmarks) and they also have different cache prices. Last, but not least, there are new emerging caching strategies that are non-standard. This adds to the confusion.
    @PMinervini@abeirami @Auss_Abbood yup that's pretty much the point of https://neuralnoise.com/2026/harness-bench-wip/?bare -- some of the entries were interview questions to potential visiting students that even early last year's frontier models could not nail; now even open-weight models solve them without issues