Researchers Note Caching Issues in Terminal Task Metrics
Conversation highlights Pareto frontier of success rate versus cost on terminal tasks and caching complications in benchmarks.
TLDR
Ahmad Beirami, a research engineer focused on RL and LLM post-training, posted on X directing attention to the Pareto frontier of success rate versus cost on terminal tasks. Leo Boytsov, a machine learning scientist who created NMSLIB, replied that counting costs is hard due to caching. Different models use caches differently, especially for agentic benchmarks, and maintain different cache prices. He added that new emerging non-standard caching strategies add further confusion.
Combined views
76K
10 Sources, first seen 35d ago