Edison Scientific Unveils BixBench3 for Biology AI Evaluation
Edison Scientific introduced the benchmark to test models on full biology paper analyses from raw data.
TLDR
Edison Scientific announced BixBench3 as its latest evaluation for language models in science. The benchmark measures whether models can generate the analyses that support entire biology papers directly from raw data. It covers 20 tasks drawn from papers and includes datasets reaching 241 GB in size. A related post from Andrew White notes that agents reproduced papers end-to-end and that frontier models scored below 50 percent on complete reproduction tasks.
Combined views
11.9K
2 Sources, first seen 35d ago
