• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

AI agent tests suggest harness edits can address stalls while fine-tuning reduces poor plans

A user describes a Stanford paper testing both approaches on travel-planning agents, separating process failures from poor plans.

1 Source, 2h ago, first seen 2h ago

TLDR

A user says a Stanford paper sorted failed travel-planning runs into process failures, such as loops and exhausted step budgets, and content failures that produced poor plans. They report that harness edits raised Qwen3.5-4B from 0.16 to 0.30 on held-out tasks and increased plan delivery from 55% to 90%, but did not reduce the share of poor plans. A LoRA fine-tuning adapter cut poor plans from 28% to 5% of Qwen3.5-9B’s held-out runs.

Combined views

—

1 Source, first seen 2h ago

— likes— comments— saves— reposts

Combined views

—

1 Source, first seen 2h ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

Rohan Paul@rohanpaul_aiNew Stanford paper finds that harness changes fix agents that loop or stall, while agents that deliver bad plans need weight training instead. An agent can be improved by editing its harness, the prompts, tools, and checks around the model, or by fine-tuning its weights. They sorted failed runs into process failures, such as loops and used-up step budgets, and content failures, where a poor plan was delivered. On a travel-planning benchmark, an LLM-driven loop rewrote the harness, and its best runs then fine-tuned the model. Harness evolution lifted Qwen3.5-4B from 0.16 to 0.30 on held-out tasks, as plan delivery rose from 55% to 90%. Harness edits never shrank the share of poor plans, but a LoRA adapter cut them from 28% to 5% of Qwen3.5-9B's held-out runs. Before improving an agent, label why its runs fail, then fix process failures in the harness and content failures in the weights.2h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    Rohan Paul@rohanpaul_aiNew Stanford paper finds that harness changes fix agents that loop or stall, while agents that deliver bad plans need weight training instead. An agent can be improved by editing its harness, the prompts, tools, and checks around the model, or by fine-tuning its weights. They sorted failed runs into process failures, such as loops and used-up step budgets, and content failures, where a poor plan was delivered. On a travel-planning benchmark, an LLM-driven loop rewrote the harness, and its best runs then fine-tuned the model. Harness evolution lifted Qwen3.5-4B from 0.16 to 0.30 on held-out tasks, as plan delivery rose from 55% to 90%. Harness edits never shrank the share of poor plans, but a LoRA adapter cut them from 28% to 5% of Qwen3.5-9B's held-out runs. Before improving an agent, label why its runs fail, then fix process failures in the harness and content failures in the weights.2h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet