AI-agent failures may show when to switch from harness tweaks to weight training
The researchers say harness changes address blocked tool calls and loops, while weight training targets poor completed plans.
TLDR
Researchers tested a self-evolving harness—the tools and runtime around a frozen model—on DeepPlanning’s multi-day travel tasks. They report that it raised Qwen3.5-4B’s held-out score from 0.16 to 0.30. Training adapters on evolved-harness trajectories added 0.13 to held-out scores under the original harness for both model sizes. The researchers propose diagnosing failures first: use harness changes for process problems and weight training when completed plans are poor.
Combined views
139
1 Source, first seen ago