医学
功率(物理)
价值(数学)
统计
显著性差异
计量经济学
数学
热力学
内科学
物理
标识
DOI:10.1213/ane.0b013e318263ca6e
摘要
To the Editor In an article discussing minimum effect size of interest, Gibbs and Weightman1 consider several scenarios based on a hypothetical study of blood pressure after tracheal intubation. The authors emphasize the importance of the power calculation when interpreting the results of clinical trials. Although I agree that the power of a study is important, I believe the authors have overstated its importance in certain situations. In one scenario (their scenario 3), the hypothetical study reports an effect that is statistically significant but smaller than the minimum effect size used in the power calculation. The authors note that there may or may not be legitimate reasons for shifting the goal posts and claiming that a smaller effect size is in fact clinically significant. However, the authors argue that such a result is not statistically robust because, based on the initial power calculations, there is a low probability that repeating the study would find a statistically significant effect. In support of this conclusion, the authors cite Goodman2 who showed that when a positive result is only moderately statistically significant, the probability is surprisingly high that a repeat of the study would fail to find a statistically significant difference. Gibbs and Weightman support this conclusion but they base their approach on the a priori power whereas Goodman used the P value and he explicitly stated that his results are “independent of the power of the original study.” It is worth considering how an observed effect can be statistically significant despite being smaller than the minimum effect size of interest used in the power calculation. This can occur when the standard deviation observed in the study is smaller than the standard deviation that was used in the power calculation, as would have been the case in Gibbs and Weightman's scenario 3. The question then arises: Which is the more valid estimate of the population standard deviation, the observed standard deviation or the a priori estimate? It could be argued that the standard deviation observed in a study sample is not always the best available estimate of the population standard deviation. The approach by Gibbs and Weightman may be preferred if the study is small and there exists a large amount of data from prior studies. However, researchers often have little data available to derive the estimated standard deviation needed in their a priori power calculations. For this reason, I believe one of the conclusions by Gibbs and Weightman is incorrect, i.e., that a positive result is necessarily less robust just because the observed effect size is smaller than the effect size used in the a priori power calculation. According to Goodman's approach, studies with the same P value are considered equally repeatable regardless of the effect size and the original power calculation. Gibbs and Weightman also discuss the interpretation of a study reporting no statistically significant difference (their scenario 1) and emphasize the importance of considering the power of the study before rejecting the possibility of a real effect. This is clearly very sound advice. However, the authors were not correct when stating that the power of the study (in this case 80%) gives the probability that a real effect actually exists (supposedly 20%). To illustrate the fallacy, imagine an absurdly underpowered study with only a handful of patients in each group and a power of 10%. If such a study fails to demonstrate an effect, it would be incorrect to conclude there is a 90% probability that a real effect exists! The value (1 − power) estimates the chance of missing an effect if it exists; it does not give the probability that an effect actually does exist. Timothy James McCulloch, MBBS, BSc(Med), FANZCA Department of Anaesthetics Royal Prince Alfred Hospital University of Sydney Camperdown NSW, Australia [email protected]
科研通智能强力驱动
Strongly Powered by AbleSci AI