Experimental questions were harder for most LLMs than conceptual ones in a bioengineering benchmark
A user shared excerpts from a BioEVAL paper reporting that several cloud-scale models exceeded 85% overall accuracy on its multiple-choice benchmark, with the leading model reaching 90%.
TLDR
A user shared excerpts from a paper describing BioEVAL as a global, multi-institutional initiative designed to assess experimental reasoning across bioengineering subfields. The quoted paper says several cloud-scale models exceeded 85% overall accuracy on its multiple-choice benchmark, with the leading model reaching 90%. It also says experimental questions were consistently harder for most models than conceptual ones, and performance varied by subfield.
Combined views
1.6K
1 Source, first seen 3h ago
