计算机科学
人工智能
判别式
稳健性(进化)
一般化
机器学习
序列学习
鉴定(生物学)
计算模型
神经肽
多任务学习
序列(生物学)
特征学习
认知
深度学习
蛋白质测序
人工神经网络
代表(政治)
作者
Jinjin Li,Xiaorui Kang,Chen Su,Hua Shi,Feifei Cui,Z. X. Zhang,Changhang Lin,Li Wei
标识
DOI:10.1021/acssynbio.6c00015
摘要
Neuropeptides are endogenous signaling molecules that regulate diverse physiological and cognitive processes. However, reliable identification from primary sequence remains challenging due to high sequence diversity, weak motif conservation, and the limited experimental annotations. Solving this challenge is crucial for elucidating the molecular structure of neural communication and accelerating the development of neuropeptide-based therapies and peptide drugs. Existing computational approaches for neuropeptide identification range from traditional machine-learning models relying on handcrafted features to deep-learning architectures that learn sequence representations. However, both types of methods struggle with the high heterogeneity and weak motif conservation of neuropeptides, resulting in limited generalization and highlighting the need for more robust predictive frameworks. To address these limitations, we propose a unified multitask neuropeptide identification framework that integrates ESM-derived protein representations, a BiLSTM encoder, and multihead self-attention to capture local and long-range sequence dependencies jointly. Within this framework, the model further leverages attention-based pooling, auxiliary knowledge distillation, and contrastive representation learning to enhance generalization and ultimately improve the accuracy and robustness of neuropeptide identification. On the independent test set, our proposed multitask learning method (NeuroPred-MTCL) demonstrates strong generalization performance, achieving an accuracy of 93.6% and an AUROC of 0.977. It further maintains a balanced trade-off between precision (92.9%) and recall (94.4%), yielding an F1-score of 0.936 and an MCC of 0.872. These results highlight the method's ability to effectively capture discriminative sequence characteristics and substantially enhance the reliability of neuropeptide identification. These results establish NeuroPred-MTCL as a robust and generalizable approach that meaningfully advances the computational identification of neuropeptides.
科研通智能强力驱动
Strongly Powered by AbleSci AI