Manual and automated agent harnesses reportedly each had an edge in drug-design tests
A blog author says manual redesign worked better for evidence and memory fixes, while automated search did better on chemistry search.
TLDR
A blog author says their team tested Qwen on three SMDD-Bench drug-design tasks, comparing a manually redesigned agent harness with one produced by automated search using Claude. An agent harness controls which tools a model can use, what evidence it sees and what it remembers. Manual redesign pulled ahead when the fix involved evidence or memory; automated search did better when the remaining challenge was searching the chemistry. The author says human judgment still mattered in diagnosing the failures.
Combined views
11.8K
6 Sources, first seen ago
