Two-sample Mendelian randomization: avoiding the downsides of a powerful, widely applicable but potentially fallible technique

孟德尔随机化 随机化 样本量测定 计算机科学 样品(材料) 医学 统计 临床试验 数学 生物 遗传学 内科学 遗传变异 化学 基因型 基因 色谱法
作者
Fernando Pires Hartwig,Neil M Davies,Gibran Hemani,George Davey Smith
出处
期刊:International Journal of Epidemiology [Oxford University Press]
卷期号:45 (6): 1717-1726 被引量:948
标识
DOI:10.1093/ije/dyx028
摘要

Mendelian randomization studies are often performed in an instrumental variables framework, using germline genetic variants as instruments for modifiable disease risk factors or exposures.1–3 Mendelian randomization analysis depends on assuming that the genetic variants: (i) are associated with the exposure (the relevance assumption); (ii) have no common cause with the outcome (the independence assumption); and (iii) have effects on the outcome that are solely mediated by the exposure (the exclusion restriction assumption).1–3 Summary data Mendelian randomization refers to methods which use summary-level instrument-exposure and instrument-outcome association results (typically, per-allele regression coefficients and standard errors) to obtain causal effect estimates. Two-sample Mendelian randomization refers to the application of Mendelian randomization methods to summary association results estimated in non-overlapping sets of individuals. These data can be obtained from the published literature, typically from summary results provided by consortia of genome-wide association studies (GWAS), or estimated directly from individual-level participant data.4 Recent examples include studies evaluating the causal effects of adiposity-related traits on risk of breast, ovarian, prostate, lung and colorectal cancers,5 of body mass index on type 2 diabetes6 and of telomere length on several health outcomes.7 As with Mendelian randomization in general, two-sample Mendelian randomization is analogous to methods originally developed in econometrics.8,9 Recent developments in two-sample Mendelian randomization are based on methods originally developed for meta-analysis.10,11 Scatter plots of all empirical Mendelian randomization studies in PubMed from 1 January 2011 to 24 October 2016. Left panel: absolute number of one-sample (dotted line) and subsample and/or two-sample Mendelian randomization studies (solid line). Right panel: proportion of subsample and/or two-sample Mendelian randomization studies (among all one-sample and subsample and/or two-sample studies). The dotted line indicates the 50% value. Even though the most natural application of summary data Mendelian randomization methods is in the two-sample setting (i.e. when instrument-exposure and instrument-outcome associations were estimated in non-overlapping sets of individuals), it is possible in principle to use summary data methods in the one-sample context – i.e. when instrument-exposure and instrument-outcome associations are estimated in the same sample. However, summary data Mendelian randomization analyses using instrument-exposure and instrument-outcome associations in the same sample or in partially overlapping samples may be prone to weak instrument bias towards the exposure-outcome estimate that would obtained using conventional methods (typically a regression of the outcome on the exposure). Therefore, using instrument-exposure and instrument-outcome associations estimated in non-overlapping samples is preferable.14 However, given that many two-sample Mendelian randomization applications use summary data from large GWAS consortia, it is possible that in many cases the instrument-exposure and instrument-outcome datasets partially overlap due to studies participating in both consortia; and detecting if this occurs and to what degree may depend on careful assessment of the description of the studies included in each consortium. The aims of this paper are to highlight the importance of harmonizing the genetic instrumental variables used for estimating the instrument-exposure and instrument-outcome associations, and to propose steps for doing this and checking the quality of the harmonization process in two-sample Mendelian randomization. In a paper recently published in the IJE, the general issue of data harmonization is discussed.15 Appropriate data harmonization is clearly essential when combining two or more independently generated datasets. This is particularly true for two-sample Mendelian randomization, because GWAS results rarely have harmonized effect (or coded) alleles. In genetic association studies, it is often assumed that genetic variants have additive (or per-allele) effects, which corresponds to coding the genotypes numerically according to the number of copies of one of the alleles. So, if a given variant was coded as AA = 0, AC = 1 and CC = 2 (i.e. according to the number of copies of the C allele), then C is termed the effect allele, and A the other (or non-coded, or baseline) allele. Table 1 provides an overview of the steps typically required to harmonize datasets of summary results of genetic associations for two-sample Mendelian randomization, based on the guidelines provided by Fortier and colleagues.15. Below, we will focus on two-sample Mendelian randomization using summary results from GWAS consortia. Overview of the data harmonization process for two-sample Mendelian randomization applications, based on the guidelines provided by Fortier and colleagues15 At minimum, the effect allele must be available in all datasets to be harmonized. Additional variables, such as the other allelea and effect allele frequency, improve the harmonization potential Missing exposure-associated variants in the variant-outcome dataset may be replaced by proxies available in the latterb but this reduces the quality of the harmonization process iii. Consider whether the populations used to generate the datasets are sufficiently similar to harmonize them At minimum, the effect allele must be available in all datasets to be harmonized. Additional variables, such as the other allelea and effect allele frequency, improve the harmonization potential Missing exposure-associated variants in the variant-outcome dataset may be replaced by proxies available in the latterb but this reduces the quality of the harmonization process iii. Consider whether the populations used to generate the datasets are sufficiently similar to harmonize them Knowing the other allele is particularly useful for harmonization of palindromic variants. Variants in high linkage disequilibrium with the index variant in the relevant ancestry group. Not having the same allele pair could be a consequence of strand orientation differences between datasets. In this case, harmonizing strand orientation will result in shared allele pairs. Alternatively, if effect allele frequencies are available, they can be used to identify if the effect allele is the major or minor allele, and such classification can be used to check allele matching. Importantly, this strategy would only be reliable if the minor allele frequency is substantially below 50%. Multiply by -1 in the case of additive effect estimates (e.g. linear regression coefficients, log(odds ratio), risk differences) or elevate to the power of -1 in the case of multiplicative effect estimates (e.g. odds ratios). 1 (or 100%) minus the effect allele frequency in the raw dataset. Overview of the data harmonization process for two-sample Mendelian randomization applications, based on the guidelines provided by Fortier and colleagues15 At minimum, the effect allele must be available in all datasets to be harmonized. Additional variables, such as the other allelea and effect allele frequency, improve the harmonization potential Missing exposure-associated variants in the variant-outcome dataset may be replaced by proxies available in the latterb but this reduces the quality of the harmonization process iii. Consider whether the populations used to generate the datasets are sufficiently similar to harmonize them At minimum, the effect allele must be available in all datasets to be harmonized. Additional variables, such as the other allelea and effect allele frequency, improve the harmonization potential Missing exposure-associated variants in the variant-outcome dataset may be replaced by proxies available in the latterb but this reduces the quality of the harmonization process iii. Consider whether the populations used to generate the datasets are sufficiently similar to harmonize them Knowing the other allele is particularly useful for harmonization of palindromic variants. Variants in high linkage disequilibrium with the index variant in the relevant ancestry group. Not having the same allele pair could be a consequence of strand orientation differences between datasets. In this case, harmonizing strand orientation will result in shared allele pairs. Alternatively, if effect allele frequencies are available, they can be used to identify if the effect allele is the major or minor allele, and such classification can be used to check allele matching. Importantly, this strategy would only be reliable if the minor allele frequency is substantially below 50%. Multiply by -1 in the case of additive effect estimates (e.g. linear regression coefficients, log(odds ratio), risk differences) or elevate to the power of -1 in the case of multiplicative effect estimates (e.g. odds ratios). 1 (or 100%) minus the effect allele frequency in the raw dataset. Based on the research question, researchers will identify (often multiple) genetic variants associated with an adequate exposure phenotype. For example, if interest lies in studying the causal effects of adiposity measures on a given outcome, then one could use as instruments genetic variants identified in GWAS of anthropometric traits such as body mass index16 or waist circumference.17 The summary-level association results for the variants that reached genome-wide significance (i.e. P < 5.0 × 10-8) – hereafter referred to as dataset 1 – are normally extracted from published papers or a web repository. Although the focus of this paper is on genetic instruments selected based on a GWAS of the exposure phenotype, it is possible that instruments are selected using other criteria. For example, genetic instruments may be selected based on results of functional studies of gene expression regulation or limited to variants located within the gene of interest. For example, the C Reactive Protein (CRP) Coronary Heart Disease (CHD) Genetics Collaboration selected four genetic variants in the CRP gene region to evaluate the causal effect of CRP on CHD risk using Mendelian randomization.18 These four variants explain 98% of the genetic variation in this locus in populations of European ancestry, and have been shown to regulate circulating CRP levels without changing the protein sequence.19 Selecting instruments this way will likely yield results that are less prone to bias due to horizontal pleiotropy compared with selecting genome-wide significant genetic instruments scattered throughout the genome.20 In this case, data from the exposure GWAS (e.g. the large CRP GWAS21) can be used to obtain precise instrument-exposure (e.g. instrument-CRP) summary association results for the previously chosen instruments (e.g. the four genetic variants in the CRP gene region). After obtaining dataset 1, the steps listed below are commonly followed: Schematic representation of chromosomes, DNA and genetic variants in a diploid cell. Ensure that all variants in dataset 1 are associated with the exposure in the same direction, typically positive (i.e. the exposure-increasing allele is the effect allele). Such standardization is important for two main reasons: (i) it is required when applying a recently-developed summary data Mendelian randomization method called MR-Egger11; and (ii) it facilitates interpretation of plots and other forms of presenting results. When a variant is not coded in the desired way, it is necessary to ‘flip’ the variant, which implies that: the effect allele becomes the other allele, and vice versa; the regression coefficient (e.g. ln(odds ratio), mean differences, etc.) must be multiplied by -1; and the effect allele frequency must be subtracted from 1. Ensuring that datasets 1 and 2 are identically coded regarding effect (e.g. exposure increasing) and other alleles. For example, if a given instrument has A as the effect allele and C as the other allele in dataset 1, then one must ensure that this same instrument is coded as having A and C as the effect and other alleles, respectively, in dataset 2. Ensuring that both datasets are coded from the same strand is very important to reduce issues with palindromic variants (as discussed in more detail below; see Figure 2). The steps described above are illustrated in Table 2. Notice that, in Table 2, genetic variants are represented by ‘rs’ followed with a number. These are called rs numbers, which uniquely identify a genetic variant and contain information such as location (i.e. chromosome and position on the chromosome), its alleles, and other useful data. Also notice that one of the genetic instruments (rs3) was missing from the outcome GWAS, so it was replaced by an LD proxy (rs5) which was available in both exposure and outcome GWAS. When it is possible to find a suitable LD proxy that is available in both exposure and outcome GWAS, it is preferable to replace the target SNP by its LD proxy in both datasets to avoid problems with phasing (discussed above). Illustration of the process of data harmonization in two-sample Mendelian randomization using fictional data. Dataset 1 corresponds to instrument-exposure associations, and dataset 2 corresponds to instrument-outcome associations. It is assumed that both datasets are coded in the forward (5’→3’) strand The genetic instrument rs3 was not available in the outcome GWAS. Therefore it was replaced by rs5 which was available in both exposure and outcome GWAS. To this end, rs5 must be in high LD with rs3 in the relevant ancestry group. LD, linkage disequilibrium; SNP, single nucleotide polymorphism; EA, effect allele; OA, other allele; EAF, effect allele frequency; NA, not available. Illustration of the process of data harmonization in two-sample Mendelian randomization using fictional data. Dataset 1 corresponds to instrument-exposure associations, and dataset 2 corresponds to instrument-outcome associations. It is assumed that both datasets are coded in the forward (5’→3’) strand The genetic instrument rs3 was not available in the outcome GWAS. Therefore it was replaced by rs5 which was available in both exposure and outcome GWAS. To this end, rs5 must be in high LD with rs3 in the relevant ancestry group. LD, linkage disequilibrium; SNP, single nucleotide polymorphism; EA, effect allele; OA, other allele; EAF, effect allele frequency; NA, not available. Depending on the situation, additional steps may be required. For example, if instruments have been identified from elsewhere (as illustrated above with the CRP example) and an exposure GWAS is only being used to obtain precise estimates of the instrument-exposure association, it is possible that the genetic instruments are missing from the exposure GWAS, and it may be necessary to replace them with LD proxies. Another step, related to but not actually part of the data harmonization process per se, is to adjust the scale of the causal effect estimate. For example, Ference and colleagues wanted to evaluate the causal effect of a 10-mmHg reduction in systolic blood pressure (exposure) on CHD (outcome).23 If, for example, instrument-exposure associations (dataset are in systolic blood pressure per of the effect allele, then it would be necessary to both the regression coefficients and standard in dataset 1 not in dataset by However, it is important to that causal effect estimates from Mendelian randomization may not the estimates from a scale because the genetic variants to differences in in the is for a genetic variants and may the exposure The would be to quality useful check is to or such as the or datasets 1 and 2, and regarding effect allele data were used to generate the it is preferable to use effect allele frequencies estimated in In Table 2, the coefficient was harmonization LD and A similar check may be to standard linear regression may be a coefficient because may be a in standard if the sample used to generate datasets 1 and 2 are quality steps include that instrument-exposure associations are all coded in the desired and that the effect and other are the same between datasets. It is possible to identify by and datasets regarding standard absolute of the regression coefficients, and minor allele (i.e. the common allele of a given genetic this may or may not be the effect be problems in each of the data harmonization These of standardization of the of the instrument-exposure of LD proxies for missing using as between the target SNP and the LD proxy when the target SNP is used in one of the datasets and the proxy is used in the other as discussed of allele harmonization between genetic variants the case of selected LD between datasets can for example, if variants are as – – and is more one variant to this and the process if harmonization is Another potential in data harmonization is strand issues – i.e. when the effect and other in dataset 1 were based on the forward (5’→3’) DNA but on the strand in dataset 2, or vice 2). Therefore, studies effects of the same SNP using for example, an SNP with in dataset 1 may be as in dataset 2. In most cases can be identified but palindromic (i.e. to that pair with each other in a DNA see Figure are to harmonize because the are the same on both These that the effect allele frequency is and that the minor allele frequency is substantially below 50% in to identify are to with palindromic that have minor allele frequencies to them by LD analyses to evaluate on Mendelian randomization or It is common in GWAS to (i.e. of genotypes of genetic variants in the using genetic datasets (e.g. and from more samples as a Therefore, in GWAS the summary results available are in to the forward strand as a consequence of to a common However, this is not a that all datasets are for analyses because studies may be to or of the same which may differences regarding strand orientation or allele To the of harmonization that we are in a single genetic instrument with A and so an for this variant can be AC or Consider that the data are coded as AA = 0, AC = 1 and CC = 2 (i.e. C is the effect allele, and A the other allele). the effect allele is often an it is possible that effect allele coding between instrument-exposure and instrument-outcome datasets. that this to genetic variant of so that the instrument-exposure association was estimated with C as the effect allele, but A was the effect allele in the instrument-outcome the allele effect estimate of the genetic instrument on the and the allele effect estimate on In this case, the allele estimates would be and (i.e. the allele estimates in the allele is not the causal effect estimated using the would be However, if the are the causal effect estimate would be This that allele in a given instrument the of its causal effect the importance of allele harmonization in summary data Mendelian randomization. recently published two-sample Mendelian randomization studies that the causal effect of CRP levels (exposure) on risk can be used to the relevance of data harmonization for two-sample Mendelian randomization. In both studies, variants were used as genetic These variants were associated with CRP < 5.0 × 10-8) in a GWAS on more from which associations were associations were obtained from the GWAS by the a method that the outcome on an additive allele in detail and colleagues an odds of per in they result as to a in CRP However, and colleagues an odds of per in when combining instruments using effects the same datasets were it was possible that were due to differences in data To we extracted regression coefficients, standard and effect of all single nucleotide identified in the CRP GWAS from the paper (i.e. dataset only the were provided in the the other were obtained using the of variants were Summary results for all genetic variants in the GWAS were from the This dataset was to only the variants associated with CRP < 5.0 × 10-8) (dataset 2). datasets were according to rs variants were identified as having the same allele between datasets. we whether the effect this case, in datasets 1 and 2. of variants effect alleles. harmonized the effect in the two the regression using the as the the to harmonized datasets an odds of per in a result that was with but not with To the for differences, we compared harmonized datasets with the datasets provided in and In analyses using the same two datasets as both and colleagues and and we that all of the that were genome-wide associated with CRP (dataset were available in the GWAS (dataset with no for so we used all and colleagues of from analyses they were genome-wide significant in the and sample they were not in the and the a that they wanted genome-wide significant variants in both and and colleagues that of the CRP genome-wide significant were not in the GWAS dataset we find all they proxies with two of but not include the other This that are used in the analyses and see available as data However, as can be in of the are all analysis and the two proxies that use are to the same associations the that they proxy both of which were included in and colleagues and and colleagues are using overlapping sets of genetic instrumental variables for analysis to the used in the and studies that such differences in variants no on the results. The coefficients and standard in and dataset 1 were This indicates that effect and other provided in were the However, when we compared dataset 2 harmonization with the between the ln(odds were and the between the standard were then compared dataset 2 with the raw dataset from the In this the between the ln(odds was that many variants in dataset 2 were coded as in the raw dataset. This was of of harmonization between datasets 1 and 2, because (as discussed it was necessary to of the variants in the dataset so that the effect allele was the allele. the between standard of a in dataset 2, the between standard be 1 if are issues with allele After we that the coefficients and standard between and the raw datasets were from the single variant This variant odds standard of and per allele in the raw and in datasets in analysis this was assumed to to the allele), not this it be due to a because the standard were all other variants in dataset 2 were coded as in the raw we to the variant the same effect and other as in the raw dataset. Therefore, were two in dataset (i) of allele harmonization between datasets 1 and 2, which in of the effect in dataset 2 not to and (ii) a regarding variant the method to as provided in an odds of per in After the variant, the odds was After harmonization exclusion of the variant with the the odds was analysis using and the odds was These results that the differences between and results were due to of allele harmonization between datasets 1 and 2 to differences in methods to the causal effect estimate or the of the issues above were in dataset 2. When applying the to the odds was When using the odds was to the result they results are shown in Table In this example, data harmonization the of the causal effect estimate from (as by and colleagues and to of per in based on Mendelian randomization analyses using the NA, not using effects using a method that the outcome on an additive allele were provided in available as data of per in based on Mendelian randomization analyses using the NA, not using effects using a method that the outcome on an additive allele were provided in available as data allele can to bias in the causal effect estimate in the Although this would not be a when the true causal effect is the estimates can be if the true causal effect is The of this bias will depend on the proportion of effect allele and in the causal effect estimate. For example, the bias may the causal effect estimate when only a genetic instruments effect allele However, when most of the instruments or the variants are the bias may be to the causal effect estimate (as in and harmonization has been only discussed in Mendelian randomization guidelines published to In this process may not be because often have to be or have to be that are To bias due to data harmonization we that researchers the harmonized datasets (as as the used in two-sample Mendelian randomization This would of the harmonization process by and having to that harmonization has been that harmonization is performed using and that the are available with the to a in that can be used to harmonize summary-level datasets of genetic associations as as data and on are available that with harmonization and other steps of two-sample Mendelian randomization such as in the which is used in the recently developed web and analysis called both the and harmonized datasets and the harmonization process with information for will likely the that in data harmonization are two-sample Mendelian randomization The process of data harmonization be included when or the analysis This would for and check harmonized in whether or not the harmonization has been harmonization is clearly essential for two-sample Mendelian randomization analyses and be performed and clearly Two-sample Mendelian randomization studies are analogous to in that they can be performed using available data and can be used to papers – in the of data or in the In the of the of the that can in this has been It is important that the of two-sample Mendelian randomization studies shown in Figure 1 not the of and and recently by The of this is by the of a paper that that combining the two could be more Mendelian randomization to a a It is of that the in the of two-sample Mendelian randomization studies could result in studies using this more common in the literature, an that be by and data are available The and the of the is by the and a for on an within the the of
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
路人发布了新的文献求助10
1秒前
大力水手完成签到,获得积分10
1秒前
ninao发布了新的文献求助10
1秒前
zdnmax发布了新的文献求助200
2秒前
CYH完成签到,获得积分10
3秒前
4秒前
4秒前
大模型应助颜老大采纳,获得10
5秒前
ZDSHI完成签到,获得积分10
6秒前
思源应助han采纳,获得10
6秒前
8秒前
8秒前
dbc1234完成签到,获得积分10
8秒前
风听完成签到 ,获得积分10
9秒前
11秒前
11秒前
yang应助壮观缘分采纳,获得10
12秒前
乐乐应助壮观缘分采纳,获得10
12秒前
大个应助夹心饼干采纳,获得10
13秒前
木子林夕完成签到,获得积分10
13秒前
淡定的梦芝完成签到,获得积分10
14秒前
15秒前
15秒前
魄罗bro发布了新的文献求助10
15秒前
16秒前
颜老大发布了新的文献求助10
16秒前
qychen完成签到,获得积分10
18秒前
NexusExplorer应助过时的亦寒采纳,获得10
19秒前
qin发布了新的文献求助10
19秒前
20秒前
Ann_完成签到,获得积分10
20秒前
情怀应助bo采纳,获得10
21秒前
ding应助体贴的小天鹅采纳,获得10
21秒前
Dlan发布了新的文献求助10
22秒前
Fuuu完成签到,获得积分10
22秒前
一一完成签到,获得积分10
22秒前
李爱国应助科研通管家采纳,获得10
23秒前
搜集达人应助科研通管家采纳,获得10
23秒前
大个应助科研通管家采纳,获得10
23秒前
SciGPT应助科研通管家采纳,获得10
23秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Health Psychology 800
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Electric machines: theory, operating applications, and controls 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
When Is Two-Stage Sample Robust Optimization Asymptotically Optimal? 500
Discerning Saints: Moralization of Intrinsic Motivation and Selective Prosociality at Work 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7593046
求助须知:如何正确求助?哪些是违规求助? 9170282
关于积分的说明 19627955
捐赠科研通 7170993
什么是DOI,文献DOI怎么找? 3267554
关于科研通互助平台的介绍 2432418
邀请新用户注册赠送积分活动 2260134