生殖系
限制
康蒂格
基因组学
生物
计算机科学
人口
计算生物学
比例(比率)
遗传力
顺序装配
启发式
体细胞
基因组
选择(遗传算法)
群体基因组学
DNA测序
全球人口
进化生物学
种系突变
数据挖掘
作者
Zhongjun Jiang,Weihua Pan,Runtian Gao,Heng Hu,Wentao Gao,Murong Zhou,Yu‐Hang Yin,Zhipeng Qian,Jin Shuilin,Guohua Wang
标识
DOI:10.1002/advs.202515308
摘要
Population genomics using short-read resequencing captures single-nucleotide polymorphisms and small insertions and deletions but struggles with structural variants, leading to a loss of heritability in genome-wide association studies. In recent years, long-read sequencing has improved pangenome construction for diverse eukaryotic species, including humans, crops, and other organisms of ecological and economic importance, addressing this issue to some extent. Sufficient-coverage high-fidelity data for population genomics is often prohibitively expensive, limiting its use in large-scale populations and broader eukaryotic species and creating an urgent need for robust low-coverage assemblies. However, current assemblers underperform in such conditions. To address this, HiFiCCL is proposed, the first assembly framework specifically designed for low-coverage high-fidelity reads, using a reference-guided, chromosome-by-chromosome assembly approach. This study demonstrates that HiFiCCL improves low-coverage assembly performance of existing assemblers and outperforms the state-of-the-art assemblers on human and plant datasets. Tested on 45 human datasets (∼5× coverage), HiFiCCL combined with hifiasm reduces the length of misassembled contigs relative to hifiasm by an average of 21.19% and up to 38.58%. These improved assemblies excel in detecting large germline structural variants, minimize inter-chromosome mis-scaffolding, and improve the detection of specific germline and tumor somatic structural variants based on the pangenome graph.
科研通智能强力驱动
Strongly Powered by AbleSci AI