200 questions reportedly reproduce an agent’s full benchmark score within 1.03 points
A post summarizing the study says adaptive testing gave the best fidelity, but the team deployed fixed subsets grouped by difficulty because they were simpler to run.
TLDR
According to a post summarizing the paper, researchers studied an analytics agent serving tens of thousands of monthly users, using 574 historical benchmark runs split by date into calibration and held-out testing periods. The post says 200 questions—38.5% of the full benchmark—reproduced the full score within 1.03 points.
An adaptive-testing method reportedly gave the best fidelity, but the team deployed simpler fixed subsets grouped by difficulty. The post says those subsets transferred to five other agent families without recalibration and stayed stable with calibration windows as short as one day.
