标杆管理
计算生物学
生物
基因组
DNA测序
全基因组测序
遗传学
癌症
计算机科学
突变
基因
业务
营销
作者
Li Tai Fang,Bin Zhu,Yongmei Zhao,Wanqiu Chen,Zhaowei Yang,Liz Kerrigan,Kurt J. Langenbach,Maryellen de Mars,Charles Lu,Kenneth B. Idler,Howard Jacob,Yuanting Zheng,Luyao Ren,Ying Yu,Erich Jaeger,Gary P. Schroth,Ogan D. Abaan,Keyur Talsania,Justin Lack,Tsai-Wei Shen
标识
DOI:10.1038/s41587-021-00993-6
摘要
The lack of samples for generating standardized DNA datasets for setting up a sequencing pipeline or benchmarking the performance of different algorithms limits the implementation and uptake of cancer genomics. Here, we describe reference call sets obtained from paired tumor–normal genomic DNA (gDNA) samples derived from a breast cancer cell line—which is highly heterogeneous, with an aneuploid genome, and enriched in somatic alterations—and a matched lymphoblastoid cell line. We partially validated both somatic mutations and germline variants in these call sets via whole-exome sequencing (WES) with different sequencing platforms and targeted sequencing with >2,000-fold coverage, spanning 82% of genomic regions with high confidence. Although the gDNA reference samples are not representative of primary cancer cells from a clinical sample, when setting up a sequencing pipeline, they not only minimize potential biases from technologies, assays and informatics but also provide a unique resource for benchmarking ‘tumor-only’ or ‘matched tumor–normal’ analyses. Tumor–normal paired DNA samples from a breast cancer cell line and a matched lymphoblastoid cell line enable calibration of clinical sequencing pipelines and benchmarking ‘tumor-only’ or ‘matched tumor–normal’ analyses.
科研通智能强力驱动
Strongly Powered by AbleSci AI