可解释性
纳米孔
纳米孔测序
人工智能
核苷酸
计算机科学
信号(编程语言)
机器学习
计算生物学
序列(生物学)
电流(流体)
DNA测序
数据挖掘
核酸序列
模式识别(心理学)
训练集
预测建模
工作(物理)
计算模型
胸腺嘧啶
生物信息学
生物系统
管道(软件)
作者
Yenan Wang,Zhixing Wu,Jia Meng
标识
DOI:10.1177/11779322251378620
摘要
Oxford nanopore sequencing enabled real-time, long-read analysis of DNA by detecting ionic current signals associated with K-mer sequences. Although many studies analyzed sequence and modification detection, our understanding of how multiple nucleotides of the K-mer sequence determine nanopore signals together is still limited. In this study, we seek to unveil the positional impact of individual nucleotide through interpretable prediction models. Multiple machine learning models were trained and optimized. To increase model interpretability and explore underlying mechanisms, the tool of SHapley Additive exPlanations was applied to make an assessment of both nucleotides and positions. Our results show that previously unseen Oxford nanopore signals were accurately predicted, and results were consistent on two different modes (R 2 = 0.9984 for 260 bps, R 2 = 0.9983 for 400 bps, R10.4 flow cell, XGBoost). Thymine bases (T) at positions 6 and 7 were the most influential, while nucleotides at positions 1, 2, 3, 4, and 9 have minimal impacts on signals. In addition, heatmap analysis toward transitions of bases revealed the impact of individual nucleotide on signal changes in a position-specific manner. Briefly, our work provided predictive and interpretable modeling of nanopore signals, concentrating on influential bases and positions among all obtainable features, which enhanced understanding of nanopore sequencing mechanisms and nucleotide/position-related signal variations.
科研通智能强力驱动
Strongly Powered by AbleSci AI