• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Rohan Paul on Harness-of-Harness for Coding Agents

    Tweet explains how persistent state aids long coding agent tasks.

    RP
    2 Sources, 25d ago, first seen 25d ago

    TLDR

    Bengaluru machine learning engineer Rohan Paul posted on X that Harness-of-Harness outperformed repeated coding-agent runs. The approach carried code, QA evidence, and plans forward across steps. Paul said long-horizon coding needs persistent project state, independent testing, and replanning from actual failures. He noted the problem that agents can forge ahead without continuity over extended projects. The post included a screenshot of an arXiv paper.

    Combined views

    8.4K

    2 Sources, first seen 25d ago

    Combined views

    8.4K

    2 Sources, first seen 25d ago

    88 likes
    88 likes
    16 comments
    53 saves
    29 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    16 comments
    53 saves
    29 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @rohanpaul_aiHarness-of-Harness beat repeated coding-agent runs by carrying code, QA evidence, and plans forward. Says long-horizon coding is not just about giving an agent more time; it needs persistent project state, independent testing, and replanning from real failures. The problem: over a long project, coding agents can forget earlier decisions, repeat work, break working features, and miss unfinished requirements. Harness-of-Harness, or HoH, fixes this by carrying the current software, test evidence, known problems, and an updated plan into the next coding run. It improved all 3 tested agent setups across 3 software benchmarks. The clearest comparison used Codex + GPT-5.5 on GameCraft-Bench: after 3 passes, HoH scored 71.52, while simply continuing the same coding agent scored 58.24. The paper also ran HoH for 70 loops and produced a playable FPS from high-level requirements. That longer test is only 1 game project and used extra tools and skills, so broader real-world generalization is still open. So for long-running coding agents, invest in persistent project state, independent QA, and evidence-driven replanning—not just more calls.

    2 Sources

    @rohanpaul_aiHarness-of-Harness beat repeated coding-agent runs by carrying code, QA evidence, and plans forward. Says long-horizon coding is not just about giving an agent more time; it needs persistent project state, independent testing, and replanning from real failures. The problem: over a long project, coding agents can forget earlier decisions, repeat work, break working features, and miss unfinished requirements. Harness-of-Harness, or HoH, fixes this by carrying the current software, test evidence, known problems, and an updated plan into the next coding run. It improved all 3 tested agent setups across 3 software benchmarks. The clearest comparison used Codex + GPT-5.5 on GameCraft-Bench: after 3 passes, HoH scored 71.52, while simply continuing the same coding agent scored 58.24. The paper also ran HoH for 70 loops and produced a playable FPS from high-level requirements. That longer test is only 1 game project and used extra tools and skills, so broader real-world generalization is still open. So for long-running coding agents, invest in persistent project state, independent QA, and evidence-driven replanning—not just more calls.