• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    OpenFold3 and the role of training data in protein–ligand structure prediction

    A user says OpenFold3 improved dramatically after training on additional private structures from pharmaceutical companies, raising concerns about claims that biological AI models generalize broadly.

    Daniel SaltzbergDS
    1 Source, 21d ago, first seen 21d ago

    TLDR

    A user argues that predicting the structures of proteins bound to other molecules is only substantially solved where training data is dense. They cite dramatic improvements in OpenFold3 after training on additional private structures from pharmaceutical companies.

    The user praises structure-prediction models but argues their success depends on both the data and how it is represented. They caution that most other biological foundation models lack structure prediction’s unusually good data curation, labels, physical constraints and well-defined representations.

    Combined views

    4.1K

    1 Source, first seen 21d ago

    Combined views

    4.1K

    1 Source, first seen 21d ago

    39 likes
    39 likes
    4 comments
    13 saves
    4 reposts
    4 comments
    13 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Daniel Saltzberg@dargasonAnother demonstration that protein-ligand structure prediction is not solved - only substantially solved where training data is dense. OpenFold3 dramatically improved when trained on additional private structures from pharma. This makes me cautious about claims of broad generalization in other bio foundation models - most do not have the unusually good data curation, labels, physical constraints, and well-defined representation of protein structure. Structure prediction models are a revelation, but their success is contingent on the data...and how that data is turned into biologically meaningful representations. I imagine the same is true in other domains.21d

    1 Source

    Daniel Saltzberg@dargasonAnother demonstration that protein-ligand structure prediction is not solved - only substantially solved where training data is dense. OpenFold3 dramatically improved when trained on additional private structures from pharma. This makes me cautious about claims of broad generalization in other bio foundation models - most do not have the unusually good data curation, labels, physical constraints, and well-defined representation of protein structure. Structure prediction models are a revelation, but their success is contingent on the data...and how that data is turned into biologically meaningful representations. I imagine the same is true in other domains.21d