Report
ActiveSaddler reportedly picks agent-training tasks by tracking unresolved failures
A post describing a Microsoft paper reports test-pass gains of 4.4 points on GAIA2 and 7.5 on Terminal-Bench 2.0.
TLDR
A post describing a Microsoft paper says ActiveSaddler chooses tasks for tuning AI agents based on unresolved failure patterns and explores unseen tasks to find new ones. Using the same optimizer, it reportedly raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. The post says reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed task order.
Combined views
2.2K
1 Source, first seen ago
