Stanford Team Launches Terminal-Bench-Science Benchmark for AI in Real Labs
Scientists from labs worldwide shared their actual research workflows to test AI agents. The result? Even top models solve just 30 percent of these authentic tasks.

A Stanford-led team released Terminal-Bench-Science v0.1, featuring 70 tasks from life, physical, Earth, math, and engineering sciences, all ported directly from researchers' computational routines with strict verification. Claude Opus 5 topped the leaderboard at 30 percent resolution, far below general benchmarks, showing how much tougher real science workflows are for AI. With 376 contributors from 22 countries, the project aims to let agents handle routine computing so scientists focus on breakthroughs; calls for version 0.2 are already open.
Combined views
4.8K
9 posts, first seen 6d ago
Stanford Team Launches Terminal-Bench-Science Benchmark for AI in Real Labs
Scientists from labs worldwide shared their actual research workflows to test AI agents. The result? Even top models solve just 30 percent of these authentic tasks.

A Stanford-led team released Terminal-Bench-Science v0.1, featuring 70 tasks from life, physical, Earth, math, and engineering sciences, all ported directly from researchers' computational routines with strict verification. Claude Opus 5 topped the leaderboard at 30 percent resolution, far below general benchmarks, showing how much tougher real science workflows are for AI. With 376 contributors from 22 countries, the project aims to let agents handle routine computing so scientists focus on breakthroughs; calls for version 0.2 are already open.