计算机科学
稳健性(进化)
财产(哲学)
图形
机器学习
人工智能
分子图
一般化
代表(政治)
理论计算机科学
训练集
钥匙(锁)
可视化
任务(项目管理)
领域知识
特征学习
数据挖掘
标记数据
数据可视化
深度学习
领域(数学分析)
外部数据表示
人工神经网络
知识图
数据建模
适用范围
图论
化学信息学
药物发现
节点(物理)
虚拟筛选
合成数据
多任务学习
作者
Van-Thinh To,Phuoc-Chung Van Nguyen,Gia-Bao Truong,Tuyet-Minh Phan,Tuyet-Minh Phan,Tieu-Long Phan,Tieu-Long Phan,Rolf Fagerberg,Peter F. Stadler,Tuyen Ngoc Truong
标识
DOI:10.1021/acs.jcim.5c01068
摘要
Molecular property prediction has become essential in accelerating advancements in drug discovery and materials science. Graph Neural Networks have recently demonstrated remarkable success in molecular representation learning; however, their broader adoption is impeded by two significant challenges: (1) data scarcity and constrained model generalization due to the expensive and time-consuming task of acquiring labeled data and (2) inadequate initial node and edge features that fail to incorporate comprehensive chemical domain knowledge, notably orbital information. To address these limitations, we introduce a Knowledge-Guided Graph (KGG) framework employing self-supervised learning to pretrain models using orbital-level features in order to mitigate reliance on extensive labeled data sets. In addition, we propose novel representations for atomic hybridization and bond types that explicitly consider orbital engagement. Our pretraining strategy is cost efficient, utilizing approximately 250,000 molecules from the ZINC15 data set, in contrast to contemporary approaches that typically require between two and ten million molecules, consequently reducing the risk of potential data contamination. Extensive evaluations on diverse downstream molecular property data sets demonstrate that our method significantly outperforms state-of-the-art baselines. Complementary analyses, including t-SNE visualizations and comparisons with traditional molecular fingerprints, further validate the effectiveness and robustness of our proposed KGG approach. The key advantages of KGG are its data efficiency and architectural versatility, driven by orbital-informed representations. By distilling essential chemical knowledge from modest corpora, it avoids extensive pretraining and excels in low-data fine-tuning, providing a robust and chemically meaningful foundation for diverse GNN architectures.
科研通智能强力驱动
Strongly Powered by AbleSci AI