概化理论
机器学习
人工智能
工作流程
一般化
计算机科学
桥(图论)
训练集
虚拟筛选
数据科学
主动学习(机器学习)
实验数据
支持向量机
作者
Marissa Dolorfino,Daniel Santos Perez,You Fu,Shu-Hang Lin,Sean McCarty,Matthew J. O’Meara,Terra Sztain
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2026-04-19
被引量:1
标识
DOI:10.64898/2026.04.18.719394
摘要
Predicting protein-ligand binding is a central challenge in computational drug discovery, and while machine learning (ML) and co-folding methods have advanced rapidly, their ability to generalize beyond training or parameterization regimes remains insufficiently understood. DNA-encoded libraries (DELs) enable ultra-large screening of billions of molecules simultaneously, providing a useful testbed for evaluating these approaches at scale. A recent NeurIPS competition revealed that even top performing ML models trained on DEL data failed at generalizing to out-of-distribution (OOD) chemical space. We investigated whether integrating structural modeling could bridge this generalization gap. We systematically assessed state-of-the-art ML, docking, and co-folding methods including Schrodinger Glide, Rosetta GALigandDock, and Boltz-2 with three biologically diverse protein targets screened against libraries containing multiple DEL synthesis formats. While ML excels in-distribution, OOD hit discrimination is dependent on both the target and ligand context, with no single method consistently dominating. These findings demonstrate that benchmark performance alone is insufficient to predict OOD performance, highlighting the need for system-dependent evaluation of binding prediction methods. We provide an open-source package for assessing protein-ligand prediction methods and analyzing high-throughput screening data: DEL-iver.
科研通智能强力驱动
Strongly Powered by AbleSci AI