Terminal-Bench-Science Benchmark Released by Stanford Team
Stanford-led effort evaluates AI agents on 70 expert-contributed research workflows across scientific domains.
TLDR
Steven Dillmann announced the release of Terminal-Bench-Science, a benchmark for AI agents on research workflows. Version 0.1 contains 70 tasks spanning life, physical, earth, mathematical, and engineering sciences. Each task was contributed by a researcher adapting their own workflows into agentic challenges on problems they care about. The project is an ongoing Stanford-led community effort built by the Terminal-Bench team together with domain experts at research institutions worldwide. Dillmann stated that it discriminates between frontier models similarly to Terminal-Bench 3.0 while lowering pass rates for every model tested on both.
Terminal-Bench-Science Benchmark Released by Stanford Team
Stanford-led effort evaluates AI agents on 70 expert-contributed research workflows across scientific domains.
