计算机科学
代表(政治)
人工智能
自然语言处理
特征(语言学)
光学(聚焦)
机器学习
环状RNA
语义学(计算机科学)
计算生物学
特征提取
鉴定(生物学)
生物医学
主题模型
数据挖掘
疾病
情报检索
生物信息学
相似性(几何)
语言模型
模式识别(心理学)
乳腺癌
作者
Mian-Shuo Lu,Lei Wang,Meng-Meng Wei,Xiaorui Su,Bo-Wei Zhao,Zhu-Hong You,De-Shuang Huang
标识
DOI:10.1109/jbhi.2026.3672861
摘要
Circular RNA (circRNA) is a kind of non-coding RNA widely present in cells. CircRNA plays a critical role in the occurrence and treatment of diseases. Unraveling the relationships between circRNAs and diseases has become a focus for diagnosis. While computational methods for predicting circRNA-disease associations (CDA) exist, they often oversimplify the representation of circRNA structures. To address this gap, we propose a novel method LMSCDA, which focuses on enhancing circRNA and disease representation by language model to predict CDAs. Specifically, we first calculate circRNA secondary structure by the chemistry principle. Then we employ a hierarchical feature extraction model to extract the circRNA structure and semantic features and amplify features by attention mechanism. Concurrently disease semantic features encoded utilize the biomedical language model. While behavioral features of circRNA and disease captured from circRNA-miRNA and circRNA-disease networks. We integrate them into comprehensive representation to predict CDAs. LMSCDA achieves an AUC of 0.9877 and an AUPR of 0.9881 in 5-fold cross-validation on the CircR2Disease dataset. Our approach yields demonstrably competitive results when evaluated against prominent existing models. Our case study on breast cancer first validated predictive accuracy of LMSCDA, with 19 of the top 20 circRNA-Breast cancer associations being confirmed by literature evidence. An analysis on independent clinical transcriptomic dataset identified highly differentially expressed circRNA by LMSCDA, pinpointing candidates for future investigation.
科研通智能强力驱动
Strongly Powered by AbleSci AI