• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Paper Claims Harness Matters More Than Model for Agents

    Rohan Paul posted about a paper arguing harness effects dominate model choice in agent benchmarks.

    RP
    1 Source, 30d ago, first seen 30d ago

    TLDR

    Rohan Paul shared a summary of research on evaluating long-horizon agents. The post states that the harness can matter more than the model itself. Benchmark scores therefore should not be compared unless the harness is disclosed or controlled. A controlled study examined several models paired with several harnesses on tasks drawn from SWE-bench Verified. The author notes that harness-induced variance proved larger than model-induced variance in the results examined.

    Combined views

    5.8K

    1 Source, first seen 30d ago

    Combined views

    5.8K

    1 Source, first seen 30d ago

    67 likes
    67 likes
    14 comments
    43 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    14 comments
    43 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @rohanpaul_aiFor long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be compared without disclosing or controlling the harness. In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks, average harness-induced variance was 7.80X model-induced variance. Keeping the model fixed, moving from the minimal to full harness changed pass@1 by 8.5 to 13.0 percentage points. Keeping the harness fixed, switching models changed it by only 2.5 to 5.0 points. The ranking itself was unstable too: 6 of 9 model-pair/harness-pair comparisons reversed under another harness. – arxiv. org/abs/2605.23950 Title: "Stop Comparing LLM Agents Without Disclosing the Harness"

    1 Source

    @rohanpaul_aiFor long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be compared without disclosing or controlling the harness. In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks, average harness-induced variance was 7.80X model-induced variance. Keeping the model fixed, moving from the minimal to full harness changed pass@1 by 8.5 to 13.0 percentage points. Keeping the harness fixed, switching models changed it by only 2.5 to 5.0 points. The ranking itself was unstable too: 6 of 9 model-pair/harness-pair comparisons reversed under another harness. – arxiv. org/abs/2605.23950 Title: "Stop Comparing LLM Agents Without Disclosing the Harness"