Paper Claims Harness Matters More Than Model for Agents
Rohan Paul posted about a paper arguing harness effects dominate model choice in agent benchmarks.
TLDR
Rohan Paul shared a summary of research on evaluating long-horizon agents. The post states that the harness can matter more than the model itself. Benchmark scores therefore should not be compared unless the harness is disclosed or controlled. A controlled study examined several models paired with several harnesses on tasks drawn from SWE-bench Verified. The author notes that harness-induced variance proved larger than model-induced variance in the results examined.
Combined views
5.8K
1 Source, first seen 30d ago