OpenFold3 and the role of training data in protein–ligand structure prediction
A user says OpenFold3 improved dramatically after training on additional private structures from pharmaceutical companies, raising concerns about claims that biological AI models generalize broadly.
TLDR
A user argues that predicting the structures of proteins bound to other molecules is only substantially solved where training data is dense. They cite dramatic improvements in OpenFold3 after training on additional private structures from pharmaceutical companies. The user praises structure-prediction models but argues their success depends on both the data and how it is represented. They caution that most other biological foundation models lack structure prediction’s unusually good data curation, labels, physical constraints and well-defined representations.
Combined views
4.1K
1 Source, first seen 1d ago
OpenFold3 and the role of training data in protein–ligand structure prediction
A user says OpenFold3 improved dramatically after training on additional private structures from pharmaceutical companies, raising concerns about claims that biological AI models generalize broadly.
TLDR
A user argues that predicting the structures of proteins bound to other molecules is only substantially solved where training data is dense. They cite dramatic improvements in OpenFold3 after training on additional private structures from pharmaceutical companies. The user praises structure-prediction models but argues their success depends on both the data and how it is represented. They caution that most other biological foundation models lack structure prediction’s unusually good data curation, labels, physical constraints and well-defined representations.