计算生物学
计算机科学
核糖核酸
DNA
错义突变
RNA结合蛋白
数据挖掘
生物
理论计算机科学
人工智能
机器学习
遗传学
突变
基因
作者
Zihao Yan,Fang Ge,Ying Zhang,Yan He,Yan Liu,Yiheng Zhu,Jiangning Song,Dong‐Jun Yu
标识
DOI:10.1021/acs.jcim.5c00756
摘要
Missense mutations in DNA- and RNA-binding proteins can disrupt vital interaction networks and frequently lead to disease. However, current methods for predicting changes in binding affinity between protein-DNA and protein-RNA remain fragmented and inefficient. In this study, we introduce MutPNI, a novel general regression model that predicts the effects of missense mutations in DNA-binding and RNA-binding proteins. To achieve this, we integrate feature embeddings from pretrained ESM-2 and ProtT5 protein language models and incorporate energy terms derived from mutant protein structures. We then train the model through a Transformer-based encoder, enabling it to attain PCC values of 0.661 and 0.701 on DNA and RNA benchmark test sets, respectively. Furthermore, our final strategy relies on the same model architecture, rather than identical parameters, to predict mutations in both protein-DNA and protein-RNA complexes, thereby highlighting shared features as well as key distinctions in the two data sets. Finally, the method's high computational efficiency allows for scaling to large biological data sets, offering a robust platform for future research. The web server and data sets for MutPNI are publicly available at https://csbioinformatics.njust.edu.cn/mutpni/ for academic use.
科研通智能强力驱动
Strongly Powered by AbleSci AI