AutoResearchExam creators say AI research agents often overfit
The creators say their benchmark gives agents 24 hours per task to experiment and improve their solutions, then checks whether those improvements hold up on data the agents never see.
TLDR
AutoResearchExam’s creators announced a benchmark covering seven research areas, including model training, data curation, AI safety and interpretability. They give agents 24 hours per task on a CPU or GPU machine and combine speed and quality in one score. The creators report that agents often overfit as they try to improve, with gains failing to fully carry over to hidden tests. In their head-to-head comparison, Astra started strongest and led for up to 19 hours, but Fable 5.1 took the top-performing spot in the final hours.
Combined views
20
1 Source, first seen 19d ago