Terminal-Bench-Science Benchmark Released by Stanford Team
Stanford-led effort evaluates AI agents on 70 expert-contributed research workflows across scientific domains.
Steven Dillmann announced the release of Terminal-Bench-Science, a benchmark for AI agents on research workflows. Version 0.1 contains 70 tasks spanning life, physical, earth, mathematical, and engineering sciences. Each task was contributed by a researcher adapting their own workflows into agentic challenges on problems they care about. The project is an ongoing Stanford-led community effort built by the Terminal-Bench team together with domain experts at research institutions worldwide. Dillmann stated that it discriminates between frontier models similarly to Terminal-Bench 3.0 while lowering pass rates for every model tested on both.
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions…

