Report
AutoSciBench aims to adapt scientific-agent benchmarks as capabilities evolve
HuggingPapers describes the framework as automatically generating benchmarks and using solver feedback to refine them.
TLDR
HuggingPapers describes AutoSciBench as a framework that automatically generates and iteratively refines benchmarks for scientific agents. It says solver feedback helps create harder tasks that adapt as agent capabilities evolve.
Combined views
1
1 Source, first seen ago
