Explainable Deep Relational Networks for Predicting Compound–Protein Affinities and Contacts

可解释性亲缘关系人工智能机器学习结合亲和力概化理论计算机科学人工神经网络化学数学立体化学生物化学统计受体

作者

Mostafa Karimi,Di Wu,Zhangyang Wang,Yang Shen

出处

期刊：Journal of Chemical Information and Modeling [American Chemical Society]
日期：2020-12-21 卷期号：61 (1): 46-66 被引量：34

链接

nih.gov arxiv.org arxiv.org nih.govdoi.org

标识

DOI：10.1021/acs.jcim.0c00866

摘要

Predicting compound-protein affinity is beneficial for accelerating drug discovery. Doing so without the often-unavailable structure data is gaining interest. However, recent progress in structure-free affinity prediction, made by machine learning, focuses on accuracy but leaves much to be desired for interpretability. Defining intermolecular contacts underlying affinities as a vehicle for interpretability; our large-scale interpretability assessment finds previously used attention mechanisms inadequate. We thus formulate a hierarchical multiobjective learning problem, where predicted contacts form the basis for predicted affinities. We solve the problem by embedding protein sequences (by hierarchical recurrent neural networks) and compound graphs (by graph neural networks) with joint attentions between protein residues and compound atoms. We further introduce three methodological advances to enhance interpretability: (1) structure-aware regularization of attentions using protein sequence-predicted solvent exposure and residue-residue contact maps; (2) supervision of attentions using known intermolecular contacts in training data; and (3) an intrinsically explainable architecture where atomic-level contacts or "relations" lead to molecular-level affinity prediction. The first two and all three advances result in DeepAffinity+ and DeepRelations, respectively. Our methods show generalizability in affinity prediction for molecules that are new and dissimilar to training examples. Moreover, they show superior interpretability compared to state-of-the-art interpretable methods: with similar or better affinity prediction, they boost the AUPRC of contact prediction by around 33-, 35-, 10-, and 9-fold for the default test, new-compound, new-protein, and both-new sets, respectively. We further demonstrate their potential utilities in contact-assisted docking, structure-free binding site prediction, and structure-activity relationship studies without docking. Our study represents the first model development and systematic model assessment dedicated to interpretable machine learning for structure-free compound-protein affinity prediction.

求助该文献

最长约 10秒，即可获得该文献文件

Explainable Deep Relational Networks for Predicting Compound–Protein Affinities and Contacts

今日热心研友