德布鲁因图
德布鲁恩序列
康蒂格
组合数学
计算机科学
算法
离散数学
数学
理论计算机科学
基因组
生物
遗传学
基因
作者
Diego Díaz-Domínguez,Taku Onodera,Simon J. Puglisi,Leena Salmela
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2022-09-07
被引量:6
标识
DOI:10.1101/2022.09.06.506758
摘要
Abstract The nodes of a de Bruijn graph (DBG) of order k correspond to the set of k -mers occurring in a set of reads and an edge is added between two nodes if there is a k − 1 length overlap between them. When using a DBG for genome assembly, the choice of k is a delicate issue: if k is too small, the DBG is tangled, making graph traversal ambiguous, whereas choosing k too large makes the DBG disconnected, resulting in more and shorter contigs. The variable order de Bruijn graph (voDBG) has been proposed as a way to avoid fixing a single value of k . A voDBG represents DBGs of all orders in a single data structure and (conceptually) adds edges between the DBGs of different orders to allow increasing and decreasing the order. Whereas for a fixed order DBG unitigs are well defined, no properly defined notion of contig or unitig exists for voDBGs. In this paper we give the first rigorous definition of contigs for voDBGs. We show that voDBG nodes, whose frequency in the input read set is in interval [ ℓ , h ] for some h and ℓ > h /2, represent an unambiguous set of linear sequences, which we call the set of ( ℓ , h )-tigs. By establishing connections between the voDBG and the suffix trie of the input reads, we give an efficient algorithm for enumerating ( ℓ , h )-tigs in a voDBG using compressed suffix trees. Our experiments on real and simulated HiFi data show a prototype implementation of our approach has a better or comparable contiguity and accuracy as compared to other DBG based assemblers.
科研通智能强力驱动
Strongly Powered by AbleSci AI