Harness-of-Harness beat repeated coding-agent runs, a paper summary says
A post summarizing the research says the approach carries code, test evidence, known problems and updated plans between runs, rather than just giving coding agents more time.
TLDR
According to the summary, Harness-of-Harness (HoH) improved all three tested agent setups across three software benchmarks. With Codex + GPT-5.5 on GameCraft-Bench, HoH scored 71.52 after three passes, versus 58.24 for simply continuing the same coding agent. The post also says a 70-loop trial produced a playable first-person shooter, but notes that this involved just one game project and extra tools and skills, leaving broader real-world generalization open.
Combined views
1K
1 Source, first seen 25d ago