生命银行
人口分层
全基因组关联研究
虚假关系
遗传关联
混淆
人口
样本量测定
关联测试
生物
计算生物学
计算机科学
统计
遗传学
医学
数学
机器学习
单核苷酸多态性
基因型
基因
环境卫生
作者
Longda Jiang,Zhili Zheng,Ting Qi,Kathryn E. Kemper,Naomi R. Wray,Peter M. Visscher,Jian Yang
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2019-04-11
被引量:65
摘要
ABSTRACT The genome-wide association study (GWAS) has been widely used as an experimental design to detect associations between genetic variants and a phenotype. Two major confounding factors, population stratification and relatedness, could potentially lead to inflated GWAS test-statistics and thereby spurious associations. Mixed linear model (MLM)-based approaches can be used to account for sample structure. However, genome-wide association (GWA) analyses in biobank samples such as the UK Biobank (UKB) often exceed the capability of most existing MLM-based tools especially if the number of traits is large. Here, we developed an MLM-based tool (called fastGWA) that controls for population stratification by principal components and relatedness by a sparse genetic relationship matrix for GWA analyses of biobank-scale data. We demonstrated by extensive simulations that fastGWA is reliable, robust and highly resource-efficient. We then applied fastGWA to 2,173 traits on 456,422 array-genotyped and imputed individuals and 2,048 traits on 46,191 whole-exome-sequenced individuals in the UKB.
科研通智能强力驱动
Strongly Powered by AbleSci AI