• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    ActiveSaddler reportedly picks agent-training tasks by tracking unresolved failures

    A post describing a Microsoft paper reports test-pass gains of 4.4 points on GAIA2 and 7.5 on Terminal-Bench 2.0.

    RP
    1 Source, 1h ago, first seen 1h ago

    TLDR

    A post describing a Microsoft paper says ActiveSaddler chooses tasks for tuning AI agents based on unresolved failure patterns and explores unseen tasks to find new ones. Using the same optimizer, it reportedly raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. The post says reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed task order.

    Combined views

    2.2K

    1 Source, first seen 1h ago

    Combined views

    2.2K

    1 Source, first seen 1h ago

    15 likes
    15 likes
    9 comments
    7 saves
    1 reposts
    9 comments
    7 saves
    1 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @rohanpaul_aiNew Microsoft paper on Automated harness optimization for agents. Most harness auto-tuners focus on how to patch prompts and tools, but which tasks produce the feedback also changes how good the final harness gets. But you will get stronger AI agents when you pick training tasks based on which failures are still unfixed, so stop feeding them a fixed task list. ActiveSaddler tracks failure patterns and works on the one most worth fixing, or tries unseen tasks to find new ones. On the same optimizer, ActiveSaddler raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. Reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed order. If you auto-tune an agent, aim your run budget at the failures that are still open.1h

    1 Source

    @rohanpaul_aiNew Microsoft paper on Automated harness optimization for agents. Most harness auto-tuners focus on how to patch prompts and tools, but which tasks produce the feedback also changes how good the final harness gets. But you will get stronger AI agents when you pick training tasks based on which failures are still unfixed, so stop feeding them a fixed task list. ActiveSaddler tracks failure patterns and works on the one most worth fixing, or tries unseen tasks to find new ones. On the same optimizer, ActiveSaddler raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. Reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed order. If you auto-tune an agent, aim your run budget at the failures that are still open.1h