Researchers Note Caching Issues in Terminal Task Metrics
Conversation highlights Pareto frontier of success rate versus cost on terminal tasks and caching complications in benchmarks.
Ahmad Beirami, a research engineer focused on RL and LLM post-training, posted on X directing attention to the Pareto frontier of success rate versus cost on terminal tasks. Leo Boytsov, a machine learning scientist who created NMSLIB, replied that counting costs is hard due to caching. Different models use caches differently, especially for agentic benchmarks, and maintain different cache prices. He added that new emerging non-standard caching strategies add further confusion.
Check out the Pareto frontier of success rate vs cost on terminal tasks and more. p.s. ox-alpha (GLM-5.3 Flash) is completely dominated by GPT-5.6 Luna and DeepSeek V4 Flash 🙃
Newly released models repeatedly appear near the top on the @terminalbench 2.1 leaderboard. Are these models actually on par with frontier models on terminal tasks? We took the top 20 models from that leaderboard and ran them on TB-fn. TB-fn reworks the same 89 tasks by adding…
Researchers Note Caching Issues in Terminal Task Metrics
Conversation highlights Pareto frontier of success rate versus cost on terminal tasks and caching complications in benchmarks.
Ahmad Beirami, a research engineer focused on RL and LLM post-training, posted on X directing attention to the Pareto frontier of success rate versus cost on terminal tasks. Leo Boytsov, a machine learning scientist who created NMSLIB, replied that counting costs is hard due to caching. Different models use caches differently, especially for agentic benchmarks, and maintain different cache prices. He added that new emerging non-standard caching strategies add further confusion.
Check out the Pareto frontier of success rate vs cost on terminal tasks and more. p.s. ox-alpha (GLM-5.3 Flash) is completely dominated by GPT-5.6 Luna and DeepSeek V4 Flash 🙃
Newly released models repeatedly appear near the top on the @terminalbench 2.1 leaderboard. Are these models actually on par with frontier models on terminal tasks? We took the top 20 models from that leaderboard and ran them on TB-fn. TB-fn reworks the same 89 tasks by adding…

