• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Harness-of-Harness Research Targets Multi-Day Autonomous Coding

    DAIR.AI post highlights paper on wrapping coding harnesses in an outer loop for continual multi-day runs.

    EL
    DA
    3 Sources, 28d ago, first seen 28d ago

    TLDR

    DAIR.AI announced research on Harness-of-Harness, an outer loop that organizes existing coding-agent harnesses into repeated planning, coding, and testing cycles. The approach aims to support multi-day autonomous software development from high-level requirements without human intervention. Authors Haoyang Yan and colleagues from Shanghai AI Lab and SJTU describe the method in a paper titled Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement. The loop balances repair against capability gains during extended runs.

    Combined views

    77.2K

    3 Sources, first seen 28d ago

    Combined views

    77.2K

    3 Sources, first seen 28d ago

    718 likes
    718 likes
    48 comments
    896 saves
    122 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    48 comments
    896 saves
    122 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @dair_ai// Harness-of-Harness // Exciting new research on coding agents that keep building for days without a human. Here is how it works: Harness-of-Harness wraps whatever coding harness you already run and organizes its executions into repeated planning, coding and testing increments. The loop balances repair against capability growth, scopes work into small verifiable steps, keeps implementation-time testing separate from independent evaluation, and constrains the outputs rather than the workflow. Across GameCraft-Bench, FrontierSWE and ProgramBench with three different harness and model pairs, it averages a 52.25 percent relative gain over the standalone harnesses after three iterations, peaking at 82.86 percent. Paper: https://arxiv.org/abs/2609.01481 Chat with Paper: https://academy.dair.ai/papers/harness-of-harness-multi-day-autonomous-software-development-with-continual-impr-2609.01481
    @omarsar0Meta harnesses are fascinating. Pay attention if you are an AI engineer. A different problem set, but ridiculously powerful if done right. I'm interested in emerging properties like better token efficiency, orchestration, routing, new test-time compute scaling...

    3 Sources

    @dair_ai// Harness-of-Harness // Exciting new research on coding agents that keep building for days without a human. Here is how it works: Harness-of-Harness wraps whatever coding harness you already run and organizes its executions into repeated planning, coding and testing increments. The loop balances repair against capability growth, scopes work into small verifiable steps, keeps implementation-time testing separate from independent evaluation, and constrains the outputs rather than the workflow. Across GameCraft-Bench, FrontierSWE and ProgramBench with three different harness and model pairs, it averages a 52.25 percent relative gain over the standalone harnesses after three iterations, peaking at 82.86 percent. Paper: https://arxiv.org/abs/2609.01481 Chat with Paper: https://academy.dair.ai/papers/harness-of-harness-multi-day-autonomous-software-development-with-continual-impr-2609.01481
    @omarsar0Meta harnesses are fascinating. Pay attention if you are an AI engineer. A different problem set, but ridiculously powerful if done right. I'm interested in emerging properties like better token efficiency, orchestration, routing, new test-time compute scaling...