纳米孔
纳米孔测序
计算机科学
顺序装配
假阳性悖论
字错误率
错误检测和纠正
算法
基因组
纳米技术
人工智能
生物
基因
材料科学
遗传学
基因表达
转录组
作者
Ying Chen,Fan Nie,Shuang Xie,Yingfeng Zheng,Qi Dai,T. A. Bray,Yao-Xin Wang,Jian-Feng Xing,Zhijian Huang,Depeng Wang,Lijuan He,Feng Luo,Jianxin Wang,Yizhi Liu,Chuan‐Le Xiao
标识
DOI:10.1038/s41467-020-20236-7
摘要
Long nanopore reads are advantageous in de novo genome assembly. However, nanopore reads usually have broad error distribution and high-error-rate subsequences. Existing error correction tools cannot correct nanopore reads efficiently and effectively. Most methods trim high-error-rate subsequences during error correction, which reduces both the length of the reads and contiguity of the final assembly. Here, we develop an error correction, and de novo assembly tool designed to overcome complex errors in nanopore reads. We propose an adaptive read selection and two-step progressive method to quickly correct nanopore reads to high accuracy. We introduce a two-stage assembler to utilize the full length of nanopore reads. Our tool achieves superior performance in both error correction and de novo assembling nanopore reads. It requires only 8122 hours to assemble a 35X coverage human genome and achieves a 2.47-fold improvement in NG50. Furthermore, our assembly of the human WERI cell line shows an NG50 of 22 Mbp. The high-quality assembly of nanopore reads can significantly reduce false positives in structure variation detection.
科研通智能强力驱动
Strongly Powered by AbleSci AI