• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Manual and automated agent harnesses reportedly each had an edge in drug-design tests

    A blog author says manual redesign worked better for evidence and memory fixes, while automated search did better on chemistry search.

    NI
    PI
    SU
    6 Sources, ,

    TLDR

    A blog author says their team tested Qwen on three SMDD-Bench drug-design tasks, comparing a manually redesigned agent harness with one produced by automated search using Claude. An agent harness controls which tools a model can use, what evidence it sees and what it remembers. Manual redesign pulled ahead when the fix involved evidence or memory; automated search did better when the remaining challenge was searching the chemistry. The author says human judgment still mattered in diagnosing the failures.

    Combined views

    11.8K

    6 Sources, first seen 3h ago

    Combined views

    11.8K

    6 Sources, first seen 3h ago

    117 likes
    3h ago
    first seen 3h ago
    117 likes
    11 comments
    49 saves
    18 reposts

    Sentiment

    Positiveβ€”β€”Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    11 comments
    49 saves
    18 reposts

    6 Sources

    @SureshRaghu07New blog: Is human taste overrated in harness engineering? An agent harness is the system around a model. It decides which tools the model can call, what evidence it sees, and what it remembers between steps. Most harnesses are still designed by human taste. Run the agent, read where it failed, change the harness, try again. That loop is increasingly being automated, with a strong model rewriting the harness itself. So we asked how far that can go. We ran Qwen on three drug design tasks from SMDD-Bench and compared a manually redesigned harness against one produced by automated harness search with Claude in the loop. No single approach won everywhere. The manual harness pulled clearly ahead when the fix was changing what the agent sees and remembers. Automated search did better once that was in place and the remaining problem was how to search the chemistry itself. The hard part was rarely implementing the fix. It was diagnosing what kind of failure we were looking at, and that is where human taste still mattered. Work with @KevinH1119568 , @aviral_kumar2 and @niloofar_mire Read more πŸ‘‡3h
    @niloofar_mireRT @SureshRaghu07: New blog: Is human taste overrated in harness engineering? An agent harness is the system around a model. It decides wh…3h
    @PrimeIntellectRL needs more long-horizon tasks beyond math and coding. @niloofar_mire's team at CMU built an environment for drug design. SMDD-Bench comprises 502 small-molecule design tasks with RDKit, ADMET-AI and Boltz-2 in the loop. The challenges go beyond chemistry: long-horizon planning, exploration, and learning from imperfect feedback are also open problems for RL/ML! SMDD-Bench is available in our Environments Hub, ready to train with prime-rl. Thank you for sharing with the community!1h

    Sentiment

    Positiveβ€”β€”Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    β€”

    Not ranked yet

    Today's Rank

    β€”

    Not ranked yet

    6 Sources

    @SureshRaghu07New blog: Is human taste overrated in harness engineering? An agent harness is the system around a model. It decides which tools the model can call, what evidence it sees, and what it remembers between steps. Most harnesses are still designed by human taste. Run the agent, read where it failed, change the harness, try again. That loop is increasingly being automated, with a strong model rewriting the harness itself. So we asked how far that can go. We ran Qwen on three drug design tasks from SMDD-Bench and compared a manually redesigned harness against one produced by automated harness search with Claude in the loop. No single approach won everywhere. The manual harness pulled clearly ahead when the fix was changing what the agent sees and remembers. Automated search did better once that was in place and the remaining problem was how to search the chemistry itself. The hard part was rarely implementing the fix. It was diagnosing what kind of failure we were looking at, and that is where human taste still mattered. Work with @KevinH1119568 , @aviral_kumar2 and @niloofar_mire Read more πŸ‘‡3h
    @niloofar_mireRT @SureshRaghu07: New blog: Is human taste overrated in harness engineering? An agent harness is the system around a model. It decides wh…3h
    @PrimeIntellectRL needs more long-horizon tasks beyond math and coding. @niloofar_mire's team at CMU built an environment for drug design. SMDD-Bench comprises 502 small-molecule design tasks with RDKit, ADMET-AI and Boltz-2 in the loop. The challenges go beyond chemistry: long-horizon planning, exploration, and learning from imperfect feedback are also open problems for RL/ML! SMDD-Bench is available in our Environments Hub, ready to train with prime-rl. Thank you for sharing with the community!1h