• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Edison Scientific Unveils BixBench3 for Biology AI Evaluation

    Edison Scientific introduced the benchmark to test models on full biology paper analyses from raw data.

    RS
    ES
    2 Sources, 35d ago, first seen 35d ago

    TLDR

    Edison Scientific announced BixBench3 as its latest evaluation for language models in science. The benchmark measures whether models can generate the analyses that support entire biology papers directly from raw data. It covers 20 tasks drawn from papers and includes datasets reaching 241 GB in size. A related post from Andrew White notes that agents reproduced papers end-to-end and that frontier models scored below 50 percent on complete reproduction tasks.

    Combined views

    11.9K

    2 Sources, first seen 35d ago

    Combined views

    11.9K

    2 Sources, first seen 35d ago

    144 likes
    144 likes
    2 comments
    87 saves
    58 reposts
    Edison Scientific, Inc

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    2 comments
    87 saves
    58 reposts

    2 Sources

    @EdisonSciIntroducing BixBench3, our latest eval for language models in science and the first eval to measure the ability of models to produce the analyses underpinning entire biology papers from raw data. BixBench3 evaluates 13 frontier models on 20 paper-derived tasks, using datasets up to 241 GB. Each task mirrors how scientists often use agents today: setting a research objective and designing a methodological plan, then delegating implementation to the agent. The best model reproduced 48% of the requested artifacts on average. Agents are starting to approach research study-scale capabilities, but there is still plenty of work to do.
    @ziv_ravidRT @andrewwhite01: Bixbench3: Can models reproduce the entire analysis from papers? We had agents repro papers end-to-end, with over 1B tok…

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @EdisonSciIntroducing BixBench3, our latest eval for language models in science and the first eval to measure the ability of models to produce the analyses underpinning entire biology papers from raw data. BixBench3 evaluates 13 frontier models on 20 paper-derived tasks, using datasets up to 241 GB. Each task mirrors how scientists often use agents today: setting a research objective and designing a methodological plan, then delegating implementation to the agent. The best model reproduced 48% of the requested artifacts on average. Agents are starting to approach research study-scale capabilities, but there is still plenty of work to do.
    @ziv_ravidRT @andrewwhite01: Bixbench3: Can models reproduce the entire analysis from papers? We had agents repro papers end-to-end, with over 1B tok…

    Related

    Kosmos reportedly drafted a drug application section in 4.2 hours

    Edison Scientific reports a first acceptable Nonclinical Overview draft for Population Health Partners, compared with roughly 100 hours of expert time.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet