Discovery Harnesses Fail to Generalize Across Model-Problem Pairs
Reactions from ranked influencers
2 postsWell at least openevolve is significant
First: discovery harnesses have a generalization problem. No fixed harness is reliably superior across model–problem pairs. Harness choice is better viewed as a model–problem-specific hyperparameter—not a universal recipe. 3/n
We found no universally best harness for automated discovery. Simpler harnesses often matched or outperformed more complex systems like OpenEvolve—and the best choice changed across model–problem pairs. To robustly evaluate this high-variance setting, we ran a controlled study using repeated budget-matched runs, strong baselines, and statistical hypothesis testing. We found: • Discovery harnesses have a generalization problem: no fixed harness is reliably superior across model–problem pairs • Greater harness complexity does not reliably improve discovery: simpler systems can match or outperform more elaborate harnesses, including variants of OpenEvolve • Choosing among harnesses online as rollouts progress alleviates this generalization problem, motivating discovery systems that adapt their harnesses during problem solving The takeaway: treat harness choice as a model–problem-specific hyperparameter, not a universal recipe—and move from fixed harnesses toward online adaptation. Paper: https://arxiv.org/abs/2607.18235 Github : https://github.com/akshat57/harness-generalization More details below 🧵 @LChoshen @GopalaSpeech @berkeley_ai
Combined views
772
2 posts, first seen 16h ago