Report
AI agent tests suggest harness edits can address stalls while fine-tuning reduces poor plans
A user describes a Stanford paper testing both approaches on travel-planning agents, separating process failures from poor plans.
TLDR
A user says a Stanford paper sorted failed travel-planning runs into process failures, such as loops and exhausted step budgets, and content failures that produced poor plans. They report that harness edits raised Qwen3.5-4B from 0.16 to 0.30 on held-out tasks and increased plan delivery from 55% to 90%, but did not reduce the share of poor plans. A LoRA fine-tuning adapter cut poor plans from 28% to 5% of Qwen3.5-9B’s held-out runs.
Combined views
—
1 Source, first seen ago
— likes— comments— saves— reposts
