摘要
In 2011, the editor of the journal Ultrasound in Obstetrics and Gynecology commissioned a systematic review showing there was a worrying lack of data from which to derive safe guidelines for the diagnosis of miscarriage and highlighted the need for larger prospective studies.1 Systematic reviews have a habit of concluding that evidence is substandard and nothing can be done pending the perfect study. However, on this occasion, the conclusions were correct and it was clear that at that time the guidelines being used to diagnose miscarriage were too open to interpretation, based on inadequate data and failed to account for both inter-observer variation and the quality of ultrasonography. For me the issue came to a head when I was chairing an early pregnancy session at the International Society for Ultrasound in Obstetrics and Gynecology meeting in Prague in 2010. I asked the audience what they used as a cut-off value for mean gestation sac diameter (MSD) to define miscarriage, and the answers ranged from a very worrying 2 to a more reassuring 20 mm. When talking to colleagues there was a sense of concern that the criteria used to diagnose miscarriage were less than watertight. From a patient's perspective if there is one decision a women would expect clinicians to get right with absolute certainty, it would be whether their baby is alive or not. This prompted us to review the evidence behind the guidelines used to make these decisions. We first turned to the American College of Radiologists (ACR) guideline, updated in 2009,2 that stated that ‘embryonic demise may be diagnosed with an embryo >5 mm without cardiac activity’. Miscarriage was also defined as an empty gestation sac measuring ≥16 mm in mean diameter. The guideline referred to two papers, one from 1988 based on 35 pregnancies with an MSD >6 mm3 and another from 1990 based on just 12 embryos with crown rump length (CRL) measurements between 4.0 and 4.9 mm.4 A further paragraph in the ACR guideline at that time in relation to pregnancy of unknown location (PUL) read: ‘These patients may also be considered to have a PUL. In this situation, the American Society of Reproductive Medicine (ASRM) advocates uterine curettage to rule out an ectopic pregnancy when the serum βhCG level is >2400 mIU/mL. This approach will undoubtedly result in the loss of some early viable intrauterine pregnancies’.2 This remarkable statement highlighted the lack of rigour associated with early pregnancy care at that time. These criteria being used to select women for curettage in this context had been highlighted previously by Condous et al.,5 who raised concerns about the risk of inadvertent termination associated with this management approach. In another study on miscarriage, Goldstein6 concluded, ‘in our hands, the absence of cardiac activity in embryos measuring 4 mm or more is reliably associated with embryonic death. In contrast, the lack of cardiac activity in embryos of 3 mm or less is non-diagnostic and may warrant follow-up study in 3–5 days’. This was certainly a definitive statement, but one based on just 22 patients. Assessing early pregnancy viability was the subject of a review published in the American Family Physician in 2009.7 This suggested an anembryonic pregnancy could be defined as the ‘presence of a gestational sac larger than 18 mm without evidence of embryonic tissues (yolk sac or embryo)’. This statement was considered to be of grade C evidence according to the SORT criteria, i.e. based on opinion and with limited evidence. The question therefore remained: for an empty gestation sac what MSD measurement should be associated with uncertain viability? One of the few prospective studies on the subject available at that time was from Elson et al.8 They examined 200 pregnancies with an empty gestation sac with an MSD <20 mm and found considerable overlap between the MSD of viable and non-viable pregnancies. They observed two pregnancies with an empty sac with an MSD between 18 and 20 mm that were subsequently shown to be viable. As the authors pointed out: ‘ultrasound is an operator-dependent method and it is conceivable that an inexperienced operator may fail to detect an embryo in a relatively large sac due to a poor examination technique’. This study was limited to pregnancies with an MSD ≤20 mm, but it rang alarm bells for those believing a cut-off value for MSD of 20 mm had a large margin of error. Bearing in mind the United States guidelines used a MSD cut-off value of ≥16 mm, the paper by Rowling et al.9 in Radiology is also relevant. In this retrospective study based on case reviews, five of 59 (8%) patients with no embryo visible and an MSD of 16 mm were subsequently recognised as viable pregnancies. What of the reproducibility of gestation sac and embryo measurements? The presence of fibroids or an axial uterus may cause technical problems when visualising an early pregnancy. However, Pexsters et al.10 showed that there is also significant clinically important variation in the accuracy of measurements for CRL and even more so for MSD. This study, though small, was prospective and conducted by two very experienced operators. These data suggested that if one operator measures MSD as 20 mm, another might measure the same gestation sac as anything between 16 and 24 mm. In this study a huge effort was made to ensure MSD was measured correctly, whilst anecdotally in our unit we still see women referred who have had a miscarriage diagnosed incorrectly on the basis of just one gestation sac diameter measurement or where the diameters have been measured incorrectly. How miscarriage was defined in the UK at that time was also unsatisfactory. The Royal College of Obstetricians and Gynaecologists (RCOG) green top guideline from 200611 defined a pregnancy of ‘uncertain viability as an: ‘intrauterine sac (MSD <20 mm) with no obvious yolk sac or fetus or fetal echo CRL <6 mm with no obvious fetal heart activity. In order to confirm or refute viability, a repeat scan at a minimal interval of 1 week is necessary’. This suggested that either an embryo with CRL ≥6 mm or empty gestation sac with MSD ≥20 mm is definitively diagnostic of miscarriage on an initial scan. The rationale for repeating a scan in 1 week was not explained. Was this on the expectation that we will see embryonic structures or that the sac would grow? Yet there are data showing the MSD can stay static for many days and still be associated with a viable embryo.12 Where are we now? In 2011, Abdallah et al.13 published a seminal paper on 1060 pregnancies of uncertain viability (PUV) showing the false positive rate for diagnosing miscarriage using a MSD of 16 mm is 4.4% and 0.5% when the MSD measures 20 mm. For an embryo CRL of 5 mm, the false positive rate is 8.3%. In an unprecedented response, the RCOG revised its guidelines within 7 days of the Abdallah paper being published, changing the cut-off values to define miscarriage to a MSD of ≥25 mm for empty gestational sacs and CRL ≥7 mm for an embryo with no heartbeat.14 This new guidance was subsequently adopted by the UK National Institute for Health and Care Excellence15 and the ACR.16 There then followed a consensus meeting of the Society of Radiologists in Ultrasound in the United States that also adopted these proposed new thresholds as well as making sensible suggestions regarding the timing of repeat scans and findings that may be worrisome but not diagnostic of miscarriage. These were reported in a review article in the New England Journal of Medicine.17 These changes were important. Recently, Hu et al.18 published an assessment of the impact of the new guidelines on the number of follow-up scans. On reviewing 1013 scans from women attending their practice with bleeding in early pregnancy, they found 125 (12%) fell into the more conservative zones defined by the new compared to the former size criteria (CRL, 5–7 mm; MSD, 16–25 mm). It has been estimated that in the UK 500,000 women attend hospital with bleeding or pain in early pregnancy. If we make a crude extrapolation, 12% is a huge number of pregnancies that were previously potentially at risk of misdiagnosis. In the asymptomatic population the situation may be worse, as the advent of highly sensitive home pregnancy tests has led to women often seeking a scan for reassurance at an early stage with a high likelihood of the scan being inconclusive.19 There are still uncertainties within national guidelines. However, more recently, in the BMJ, Preisler et al.20 reported on 2854 PUV. This study has confirmed that the new cut-off values for MSD and CRL used to define miscarriage do so with high specificity and narrow confidence intervals. This paper also made important observations on the timing of repeat scans, what structures should be seen on these scans and the impact of gestational age. The authors make proposals to further refine the diagnostic criteria for miscarriage, the most important being the need to repeat scans after at least 14 days in the event of relatively small empty gestational sacs (MSD <12 mm). A search on PubMed finds multiple papers relating to possible causes and treatment of recurrent miscarriage, a problem affecting a very small percentage of our early pregnancy population. Yet, prior to 2011, we had only a few small studies on how to diagnose miscarriage and a situation where MSD measurements used to define miscarriage ranged from 15 to 25 mm even in published guidelines. A woman could fly from New York to London and be told she had a miscarriage in one country and a PUV in the other. This is exactly what the Institute of Medicine was referring to in its recent report21 when it stated that diagnostic error is a ‘moral, professional, and public health imperative’, highlighting that diagnostic errors are one of the most common issues affecting patient safety. The tendency to prioritise the funding and publication of randomised trials or trials involving interventions needs to be reviewed. In all fields a little more attention should be given to getting the diagnosis right first, with good quality diagnostic studies being ranked appropriately. As clinicians with an interest in ultrasonography we all understand the need for accurate diagnosis. We need to push the agenda that diagnostic accuracy trials are considered important. An example is seen in the UK where in 2009 only 35% of women with ovarian cancer had their surgery carried out by an appropriately trained gynaecological oncologist.22 Conversely, many benign masses probably have surgery inappropriately. An accurate initial diagnosis in these cases would surely improve this situation, and we have diagnostic tests available that are able to do this.23, 24 These problems do not only concern clinicians. Patient expectation needs to be managed, and patient groups have a role to play by educating women about diagnostic performance and that a result is not always possible on the basis of one scan or blood test. We have similar issues to early pregnancy in relation to the need to repeat ultrasound scans to rule out physiological cysts of the ovary, allow menstruation to occur to better evaluate the endometrium, or simply because the views obtained were suboptimal on a given day. The issue of diagnostic test performance matters. It is axiomatic when diagnosing miscarriage that one mistake is too many for any couple looking forward to their pregnancy. It is now over 20 years since the landmark Cardiff enquiry in the UK that led to a common sense report by Hately et al.25 that emphasised the Hippocratic oath ‘to do no harm’. This report should be compulsory reading for all healthcare practitioners who work in the care of women in early pregnancy. The authors were clear that they recorded the events in Cardiff ‘not to cast blame but to show what may happen in any busy ultrasound practice unless proper protocols and precautions are established’. For me, the key phrase in the Cardiff report is ‘the death of an early pregnancy should be regarded as of equal significance to that occurring at a later stage’. This sentence should be printed on the ultrasound machine in any unit that examines women in early pregnancy. Tom Bourne is supported by the National Institute for Health Research (NIHR) Biomedical Research Centre based at Imperial College Healthcare NHS Trust and Imperial College London. The views expressed are those of the author(s) and not necessarily those of the NHS, the NIHR or the Department of Health.