摘要
We noted with interest the recent systematic review of reliability studies in obstetrics and gynecology published in this Journal1 and agree that this is an area worthy of clarification. Unfortunately, the authors have chosen to ‘judge’ study quality based predominantly upon self-citation of their own previous Correspondence2, which proposed acceptable levels for repeatability measures for ultrasound publications without factual substantiation. The authors reference a previous GRRAS paper (though probably intended to cite the version published in an epidemiology journal)3. This itself was based upon informed opinion rather than factual content, did not concern ultrasound (or more specifically Doppler) repeatability, and did not in fact contain any such proscriptive statements for ‘acceptable’ intraclass correlation coefficients (ICCs)3. Physiological variation, including beat-to-beat variation, is well recognized in Doppler ultrasound, as can be seen at routine examination by turning on the continuous measurement function of pulsed-wave Doppler. This variation generates a ‘random error’ that would be additive to human measurement error in measuring an ICC for human repeatability. If one, for example, assumed this random error to be in the order of ±10%, repeatability studies based upon multiple Doppler waveforms would have an ICC that was physiologically restricted to a maximum of 0.90. Researchers could potentially achieve higher ICCs by simply repeating measurements of a single waveform but this would not be reflective of fetal physiology. Our discipline is, in fact, full of examples of situations in which clinical practice has been influenced by Doppler studies that would have failed according to the authors' seemingly unattainable criteria. Determination of middle cerebral artery peak systolic velocity (MCA-PSV) for detection of fetal anemia, for example, had only ‘poor’ to ‘moderate’ ICCs of 0.82–0.95 amongst nine observers under real-world fetal medicine unit conditions4, while the authors' own work demonstrating poor reliability of umbilical artery resistance indices5 does not negate the fact that monitoring of umbilical artery Doppler decreases perinatal mortality in high-risk pregnancies6. We feel that the unilateral setting of such standards is destructive to the performance of accurate research, and does little to facilitate honesty of presentation of research findings. We agree that a definition of ‘acceptable, good or excellent’ repeatability metrics for Doppler repeatability studies is needed, though feel this needs to be a consensus view and to incorporate known physiological variation. In designing such a consensus approach, we propose that distinction be made between fixed measurements, such as fetal biometry, and those features that are known to exhibit physiological variation, such as indices of cardiac function or Doppler blood flow indices. Additionally, there must be recognition that near-perfect reliability is of no use if proposed new techniques do not reflect (patho)physiology adequately. Conversely, when techniques, such as MCA-PSV, arise that genuinely allow this ultrasonic distinction, reliability requirements need not be as stringent for the technique to translate appropriately to clinical practice. A. Welsh*†‡ and A. Henry§ †Royal Hospital for Women, Department of Maternal-Fetal Medicine Randwick, Sydney, New South Wales, Australia; ‡University of New South Wales, Division of Women & Children's Health, Randwick, Sydney, New South Wales, Australia; §Department of Obstetrics & Gynaecology, St George Hospital, Kogarah, New South Wales, Australia *Correspondence. (e-mail: [email protected])