How much power is enough? Against the development of an arbitrary convention for statistical power calculations

作者
Julian Di Stefano
出处
期刊:Functional Ecology [Wiley]
卷期号:17 (5): 707-709 被引量:159
标识
DOI:10.1046/j.1365-2435.2003.00782.x
摘要

Due largely to the work of Jacob Cohen (1962, 1988), statistical power analysis has become a widely known technique in many fields of experimental science. Promotion of power analysis in ecology (Toft & Shea 1983; Peterman 1990a,b; Fairweather 1991) has led to its increasing use over the last few decades. This trend is encouraging as power analysis is a useful tool in the planning phase of ecological experiments and can also be used after data analysis to improve the interpretation of non-significant results (Quinn & Keough 2002). Statistical power = 1 –β where β is the Type II error rate. Power can generally be described by the equation Power ∝ (ES × α × √n)/σ where ES is the effect size, α is the Type I error rate, n is the sample size and σ is the population standard deviation. A Type I error is the probability of erroneously rejecting the null hypothesis, while a Type II error is the probability of erroneously failing to reject the null hypothesis. Equation 1 (its specific form depends on the statistical model being used) is usually solved for power, sample size or effect size depending on the objectives of the analysis (Fairweather 1991), and can be implemented either before or after data are collected. Whichever way power analysis is used, researchers must often make a decision about what constitutes an acceptable level of power. This is almost always difficult as there is no simple way to decide how much power is enough. Unfortunately, the definition of adequate statistical power in the ecological literature often appears arbitrary with minimal attention to the context within which individual studies are conducted. Authors frequently state that statistical power is adequate if its value is 0·80 or above (Steidl, Hayes & Schauber 1997; Lougheed, Breault & Lank 1999; Manolis, Andersen & Cuthbert 2000; Strehlow et al. 2002), and a convention is fast developing where ‘significance and power levels are set … at 0·05 and 0·80 … respectively’ (Walsh et al. 1999). This practice is hereafter referred to as the five-eighty convention. When the five-eighty convention is used, the probabilities of making Type I and Type II errors are 5% and 20%, respectively. Implicitly, this means that the cost of making a Type I error is considered four times more important than the cost of making a Type II error (Cohen 1988). While in some situations this may be the case, it certainly cannot be assumed. For example, in the fields of impact assessment and conservation biology, making a Type II error will often be more costly than making a Type I error (Peterman 1990b; Taylor & Gerrodette 1993). However, the specific scientific, economic and socio-political context within which research is conducted will influence the relative importance of each type of error (Di Stefano 2001), thus designating research fields in which Type I or Type II errors are more important is problematic. The relative costs of statistical errors need to be considered each time that an experiment or monitoring program is planned. The need to identify the relative costs of Type I and Type II errors when deciding on an acceptable level of statistical power has been discussed at some length in the literature (Toft & Shea 1983; Peterman 1990a; Peterman & Mgonigle 1992; Mapstone 1995; Keough & Mapstone 1997; Downes et al. 2002). Although evaluating the relative costs of Type I and Type II errors is complex and may be influenced by the perspective of the decision maker, the process brings the functional consequences of statistical errors to the fore and thus promotes rational consideration of the scientific, economic or socio-political issues that may be at stake. For example, consider an experiment designed to test the ecological impact of a toxin in a waterway. In this case the cost of making a Type II error (concluding there is no toxic effect when there is one) is arguably greater than the cost of making a Type I error (concluding that there is a toxic effect when one does not exist). A Type I error may result in unnecessary clean up efforts, or the implementation of unjustified fines or other penalties. A Type II error, however, would result in management inaction and (depending on the type and quantity of the toxin) potentially serious environmental damage. Peterman (1990b) provides an example from the field of fisheries ecology where making a Type II error led to management inaction and a subsequent decline in fish stocks. Alternatively, if researchers were interested in testing whether a new highly mechanised timber-harvesting technique resulted in better wildlife habitat than a conventional labour-intensive technique, the cost of making a Type I error (concluding that the new technique is better when it is not) may be greater. If the implementation of the new technique resulted in large technology changeover costs and job losses, making a Type I error could place substantial economic pressure on the local human population. A Type II error (concluding that the new technique made no difference when in fact it results in better habitat) may be less important if there were no pressing reason to improve wildlife habitat in the area. In other situations, the costs of Type I and Type II errors may be the same. Clearly, appropriate values for α and β (and hence power) should not be fixed, but should vary depending on circumstances specific to each experiment or monitoring program. An example of this approach involves defining an acceptable ratio of α:β in the planning phase of an experiment, deciding on ideal values for α and power and then computing the sample size required to achieve these values. If achieving this sample size falls beyond the project budget (a common occurrence when experimental units are large or data are inherently variable) new values for α and power can be assigned, but the initial α:β ratio does not change; maintenance of this ratio is important as it preserves the relative cost of Type I and Type II errors throughout the process of sample size determination (Mapstone 1995; Keough & Mapstone 1997; Downes et al. 2002). Following this process not only facilitates rational determination of error rates and power but enables researchers to assess the feasibility of their initial objectives. For cases where power is low, raising α within acceptable limits or redesigning the study to increase power (e.g. Foster 2001) will improve the clarity, precision and usefulness of statistical outputs and thus may even increase the likelihood of publication. Research recently published in Functional Ecology serves to illustrate the inappropriate use of the five-eighty convention. Perkins & Speakman (2001) investigated whether measuring the abundance of 13C in the breath of laboratory mice could be used to detect differences in their diet. One of their experiments used anova to detect differences in 13C abundance between three groups of 10 mice fed diets of wheat, maize and mealworm, respectively. Although the anova detected differences between maize and wheat diets and maize and mealworm diets, differences between the wheat and mealworm diets were not detected due to high variance and a small effect size. Using the five-eighty convention, Perkins and Speakman predicted that if the experiment were conducted again, 41 mice per group would be required to establish a statistically significant difference between the wheat and mealworm diets on the basis of 13C abundance. This calculation was informative because it suggested that, in a future experiment, large numbers of sample animals would be needed to differentiate between diet types if the 13C signature of different diets was similar. However, the use of the five-eighty convention implied that the cost of a Type I error was four times more important than the cost of a Type II error when this was clearly not the case. In the context of a future experiment where diets are unknown, a Type I error would mean detecting a difference between diets when diets were the same, and a Type II error would mean concluding that diets were the same when in fact they were different. Given no further information, it is reasonable to assume that the cost of Type I and Type II errors in an experiment of this nature would be approximately equal. Thus, if α was set at 0·05, power should have been 0·95. If these values were used in the power calculation the number of mice per group rises from 41 to 67. Although the general conclusion (that a large sample of mice would be needed to establish a statistically significant difference in 13C abundance) does not change, a sample size of 67 has a logical basis while a sample size of 41 does not. Theoretically, other combinations of α and power would also be acceptable as long as the α:β ratio remained at 1:1 – it would simply depend on the level of error probability that researchers (and perhaps other stakeholders) were prepared to accept. Ironically, the increasing use of the five-eighty convention may be due to Jacob Cohen, or rather to the misrepresentation of comments made in his well known power analysis book (Cohen 1988). Cohen proposed the use of the five-eighty convention (Cohen 1988, p. 56), and this has sometimes been seen to legitimise its use in ecological studies (Lougheed et al. 1999; Walsh et al. 1999). Cohen's comments, however, are made with reference to the field of behavioural psychology where (he suggests) the cost of Type I errors is usually greater than the cost of Type II errors. This is frequently not the case in ecological studies. In addition, Cohen suggests the use of the five-eighty convention only when researchers have no other basis for setting the desired level of power. In almost every case, a rational basis for determining an adequate level of power can be gained by considering the relative costs of Type I and Type II errors. If the results of ecological studies are to be considered rational and logical, the process of data analysis and interpretation itself needs to be as rational and logical as possible. Reverting to baseless conventions does not help to achieve this end. Using the five-eighty convention for power analysis may be appealing because it suggests objectivity and removes the need to make difficult decisions about the relative cost of Type I and Type II errors. Nevertheless, values of α and power have important implications for the interpretation of power analyses and should not be set arbitrarily. Advice by Mapstone (1995), Keough & Mapstone (1997) and others suggesting a logical process for determining these values in a priori power calculations is sound and should be followed. Similar logical processes should be applied when calculating power retrospectively. Thanks to Lauren Bennett, Sabine Kasel, Jan Carey and Alan York for commenting on early drafts, and to Tom McKenzie for inspiration. Thanks also to Sarah Perkins and John Speakman who facilitated the sample size recalculation using their original data, and to John Hutchinson and another anonymous referee for making valuable suggestions.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
前男友发布了新的文献求助10
刚刚
小二郎应助doppelganger采纳,获得10
1秒前
无极微光应助无辜汉堡采纳,获得20
1秒前
jiang发布了新的文献求助10
1秒前
An发布了新的文献求助10
1秒前
1秒前
2秒前
2秒前
owldan完成签到,获得积分10
3秒前
3秒前
molihuakai应助felix采纳,获得10
3秒前
馒头发布了新的文献求助10
4秒前
香蕉觅云应助雨雨采纳,获得10
4秒前
4秒前
4秒前
打打应助无限的妖妖采纳,获得10
6秒前
7秒前
7秒前
想要毕业发布了新的文献求助10
8秒前
8秒前
8秒前
Wang发布了新的文献求助10
8秒前
8秒前
1eader1发布了新的文献求助10
9秒前
zyy完成签到,获得积分10
9秒前
馒头完成签到,获得积分10
9秒前
9秒前
撒大苏打完成签到,获得积分10
9秒前
10秒前
慕青应助飞快的千万采纳,获得10
10秒前
冷艳的璎发布了新的文献求助30
10秒前
10秒前
11秒前
魏佳奇发布了新的文献求助10
11秒前
11秒前
斯文败类应助晴宝采纳,获得10
11秒前
HARDCARBON完成签到,获得积分10
12秒前
hanliulaixi发布了新的文献求助10
12秒前
clwh2006完成签到,获得积分10
14秒前
漠之梦发布了新的文献求助10
14秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Römisch-Germanische Forschungen 1000
APA handbook of comparative psychology: Basic concepts, methods, neural substrate, and behavior 1000
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7609725
求助须知:如何正确求助?哪些是违规求助? 9185330
关于积分的说明 19676499
捐赠科研通 7183436
什么是DOI,文献DOI怎么找? 3270328
关于科研通互助平台的介绍 2433976
邀请新用户注册赠送积分活动 2264807