Pan-genome Search and Storage

作者
Guillaume Holley
出处
期刊:Bielefeld University - PUB – Publications at Bielefeld University
摘要

High Throughput Sequencing (HTS) technologies are constantly improving and making genome sequencing more affordable. However, HTS sequencers can only produce short overlapping genome fragments that are erroneous and cover the sequenced genomes unevenly. These genome fragments are assembled based on their overlaps to produce larger contiguous sequences. Since de novo genome assembly is computationally intensive, some species have a reference genome used as a guide for assembling genome fragments from the same species or as a basis for comparative genomics methods. Yet, assembling a genome is an error-prone process depending on the quality of the sequencing data and the heuristics used during the assembly. Furthermore, analyses based on a reference are biased towards the reference. Finally, a single reference cannot reflect the dynamics and diversity of a population of genomes. Overcoming these issues requires to move away from the single-genome reference-centric paradigm and take advantage of the multiple sequenced genomes available for each species. For this purpose, pan-genomes were introduced as sets of genomes from different strains of the same species. A pan-genome is represented by a multi-genome index exploiting the similarity and redundancy of the genomes it contains. Still, pan-genomes are more difficult to analyze than single genomes because of the large amount of data to be stored and indexed.<br /><br />\n\nCurrent data structures for pan-genome indexing do not fulfill all requirements for pan-genome analysis. Indeed, these data structures are often immutable while the size of a pan-genome grows constantly with newly sequenced genomes. Frequently, these data structures consider only assemblies as input, while unassembled genome fragments abound in databases. Also, indexing variants and similarities between the genomes of a pan-genome usually requires time and memory consuming algorithms such as sequence alignments. Sometimes, pan-genome analysis tools just assume variants and similarities are provided as input.\nWhile data structures already exist for pan-genome indexing, no solution is currently proposed for genome fragment compression in a pan-genome context. Indeed, it is often of interest to transmit and store all genome fragments of a pan-genome. However, HTS-specific compression tools are not dynamic and cannot update a compressed archive of genome fragments with new fragments of a genome without decompression. Hence, those tools are poorly adapted to the transmission and storage of genome fragments in a pan-genome context.<br /><br />\n\nIn this thesis, we aim to provide scalable solutions for pan-genome indexing and storage. We first address the problem of pan-genome indexing by proposing a new alignment-free, reference-free and incremental data structure that considers genome fragments as well as assemblies in input: the Bloom Filter Trie (BFT). The BFT is a tree data structure representing a colored de Bruijn graph in which k-mers, words of length k from the input genomes, are associated with sets of colors representing the genomes in which they occur. The BFT makes extensive use of Bloom filters to navigate in the tree and optimize the graph traversal. A "bursting" method is employed to perform an efficient path and level compaction of the tree. We show that the BFT outperforms a data structure that has similar features but is based on an approximation of the set of indexed k-mers.<br /><br />\n\nSecondly, we address the problem of genome fragments compression in a pan-genome context by proposing a new abstract data structure, the guided de Bruijn graph. It augments the de Bruijn graph with k-mer partitions such that the graph traversal is guided to reconstruct exactly the genome fragments when decompressing. Different techniques are proposed to optimize the storage of fragments in the graph and the partition encoding. We show that the BFT described previously has all features required to index a guided de Bruijn graph and is used in the implementation of our compression method named DARRC. The evaluation of DARRC on a large pan-genome dataset compared to state-of-the-art HTS-specific and general purpose compression tools shows a 30% compression ratio improvement over the second best performing tool of this evaluation.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
1秒前
科研通AI6.2应助xiaohululu采纳,获得100
1秒前
peppa发布了新的文献求助10
1秒前
芒go完成签到,获得积分10
1秒前
Basang发布了新的文献求助10
2秒前
梨白发布了新的文献求助10
2秒前
2秒前
Enma完成签到,获得积分10
3秒前
4秒前
无花果应助微末采纳,获得10
4秒前
神勇难胜发布了新的文献求助10
5秒前
5秒前
顾矜应助Yiming采纳,获得10
5秒前
6秒前
fyz完成签到,获得积分20
6秒前
情怀应助王王采纳,获得10
6秒前
Jenna发布了新的文献求助10
7秒前
7秒前
星落枝头完成签到,获得积分10
7秒前
8秒前
卷清发布了新的文献求助10
10秒前
Andrew发布了新的文献求助10
10秒前
星落枝头发布了新的文献求助10
11秒前
11秒前
11秒前
12秒前
神勇难胜完成签到,获得积分20
12秒前
13秒前
桐桐应助小茴香豆采纳,获得10
13秒前
13秒前
李爱国应助111采纳,获得10
15秒前
研友_VZG7GZ应助卷清采纳,获得10
15秒前
16秒前
fujunhao发布了新的文献求助10
16秒前
16秒前
领导范儿应助hfbbaby采纳,获得10
16秒前
cui完成签到,获得积分10
16秒前
Yiming发布了新的文献求助10
17秒前
阳光问安完成签到 ,获得积分0
18秒前
高分求助中
Les chinois de jakarta: temples et vie collective 1000
Autoparametric Resonance in Mechanical Systems 1000
Social Psychology 800
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 800
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7648891
求助须知:如何正确求助?哪些是违规求助? 9221448
关于积分的说明 19794900
捐赠科研通 7214669
什么是DOI,文献DOI怎么找? 3277947
关于科研通互助平台的介绍 2438966
邀请新用户注册赠送积分活动 2276294