Edison Scientific Unveils BixBench3 for Biology AI Evaluation
Edison Scientific introduced the benchmark to test models on full biology paper analyses from raw data.
Edison Scientific announced BixBench3 as its latest evaluation for language models in science. The benchmark measures whether models can generate the analyses that support entire biology papers directly from raw data. It covers 20 tasks drawn from papers and includes datasets reaching 241 GB in size. A related post from Andrew White notes that agents reproduced papers end-to-end and that frontier models scored below 50 percent on complete reproduction tasks.
Introducing BixBench3, our latest eval for language models in science and the first eval to measure the ability of models to produce the analyses underpinning entire biology papers from raw data. BixBench3 evaluates 13 frontier models on 20 paper-derived tasks, using datasets…
