计算机科学
化学
聚类分析
机器学习
图形
人工智能
虚拟筛选
水准点(测量)
数据挖掘
分类器(UML)
一般化
人工神经网络
深层神经网络
班级(哲学)
深度学习
化学信息学
大数据
作者
Haihan Liu,Jiaqi Lin,Ying Fan,Jintong Du,Xinying Yang,Hao Fang,Xuben Hou
标识
DOI:10.1021/acs.jcim.6c02160
摘要
Abstract The performance of target-specific, ligand-based virtual screening models is strongly influenced by dataset characteristics, including data availability, class imbalance, and evaluation strategies. In this work, we perform a systematic evaluation of graph neural networks (GNNs) using a ChEMBL-derived dataset spanning 5,368 targets and 1.59 million activity records, capturing the long-tailed distributions and target-specific imbalances commonly observed in pharmaceutical data. Through a systematic evaluation of multiple GNN architectures, we identify guidelines for model selection: while the Graph Isomorphism Network (GIN) consistently outperforms others on datasets with >100 samples (a mean ROC–AUC up to 0.94), simpler architectures are more robust under extreme data scarcity. Critically, our comparative analysis of splitting strategies reveals that random sampling yields artificially optimistic performance due to structural overlaps, whereas similarity-aware clustering exposes a substantial generalization gap (AUC drop > 0.3), cautioning against prevailing evaluation practices. We further demonstrate that multi-task learning serves as an effective remedy for small-target instability, providing significant and consistent performance gains. To underscore its translational value, we deploy this comprehensive framework in a virtual screening campaign against Mcl-1, yielding a chemically optimized lead, C4 (Ki = 0.58 μM), with verified cellular efficacy. Our findings highlight the importance of task-aware benchmark design and offer a practical strategy for reliable GNN application in drug discovery.
科研通智能强力驱动
Strongly Powered by AbleSci AI