Models Show Pass@3 Drops on Benchmark Variants
Fidian tested top models on a variant of Terminal-Bench using the same 89 tasks.
TLDR
Ahmad Beirami, a research engineer focused on RL and LLM post-training, posted that some models exhibit significant pass@3 drops when tasks are replaced with variants. The claim references tests by Fidian that ran the top 20 models from Artificial Analysis's Terminal-Bench 2.1 leaderboard on both the original benchmark and TB-fn. TB-fn is Fidian's variant built from the same 89 tasks as TB-2.1. The post calls the drops suspicious.
Combined views
1 Source, first seen 32d ago