Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study

肺炎 横断面研究 医学 一般化 射线照相术 人工智能 内科学 放射科 病理 计算机科学 数学 数学分析
作者
John R. Zech,Marcus A. Badgeley,Manway Liu,Anthony Costa,J. Titano,Eric K. Oermann
出处
期刊:PLOS Medicine [Public Library of Science]
卷期号:15 (11): e1002683-e1002683 被引量:1497
标识
DOI:10.1371/journal.pmed.1002683
摘要

BACKGROUND: There is interest in using convolutional neural networks (CNNs) to analyze medical imaging to provide computer-aided diagnosis (CAD). Recent work has suggested that image classification CNNs may not generalize to new data as well as previously believed. We assessed how well CNNs generalized across three hospital systems for a simulated pneumonia screening task. METHODS AND FINDINGS: A cross-sectional design with multiple model training cohorts was used to evaluate model generalizability to external sites using split-sample validation. A total of 158,323 chest radiographs were drawn from three institutions: National Institutes of Health Clinical Center (NIH; 112,120 from 30,805 patients), Mount Sinai Hospital (MSH; 42,396 from 12,904 patients), and Indiana University Network for Patient Care (IU; 3,807 from 3,683 patients). These patient populations had an age mean (SD) of 46.9 years (16.6), 63.2 years (16.5), and 49.6 years (17) with a female percentage of 43.5%, 44.8%, and 57.3%, respectively. We assessed individual models using the area under the receiver operating characteristic curve (AUC) for radiographic findings consistent with pneumonia and compared performance on different test sets with DeLong's test. The prevalence of pneumonia was high enough at MSH (34.2%) relative to NIH and IU (1.2% and 1.0%) that merely sorting by hospital system achieved an AUC of 0.861 (95% CI 0.855-0.866) on the joint MSH-NIH dataset. Models trained on data from either NIH or MSH had equivalent performance on IU (P values 0.580 and 0.273, respectively) and inferior performance on data from each other relative to an internal test set (i.e., new data from within the hospital system used for training data; P values both <0.001). The highest internal performance was achieved by combining training and test data from MSH and NIH (AUC 0.931, 95% CI 0.927-0.936), but this model demonstrated significantly lower external performance at IU (AUC 0.815, 95% CI 0.745-0.885, P = 0.001). To test the effect of pooling data from sites with disparate pneumonia prevalence, we used stratified subsampling to generate MSH-NIH cohorts that only differed in disease prevalence between training data sites. When both training data sites had the same pneumonia prevalence, the model performed consistently on external IU data (P = 0.88). When a 10-fold difference in pneumonia rate was introduced between sites, internal test performance improved compared to the balanced model (10× MSH risk P < 0.001; 10× NIH P = 0.002), but this outperformance failed to generalize to IU (MSH 10× P < 0.001; NIH 10× P = 0.027). CNNs were able to directly detect hospital system of a radiograph for 99.95% NIH (22,050/22,062) and 99.98% MSH (8,386/8,388) radiographs. The primary limitation of our approach and the available public data is that we cannot fully assess what other factors might be contributing to hospital system-specific biases. CONCLUSION: Pneumonia-screening CNNs achieved better internal than external performance in 3 out of 5 natural comparisons. When models were trained on pooled data from sites with different pneumonia prevalence, they performed better on new pooled data from these sites but not on external data. CNNs robustly identified hospital system and department within a hospital, which can have large differences in disease burden and may confound predictions.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
鲜艳的乐珍完成签到,获得积分10
刚刚
It完成签到 ,获得积分10
刚刚
zzzzzyq完成签到 ,获得积分10
2秒前
韩佩阳完成签到,获得积分10
4秒前
百里守约完成签到 ,获得积分10
5秒前
Sure完成签到 ,获得积分10
5秒前
时间尘埃完成签到,获得积分10
8秒前
ikun完成签到 ,获得积分10
11秒前
是阿龙呀完成签到 ,获得积分10
11秒前
橙橙完成签到 ,获得积分10
11秒前
沉静的书南完成签到,获得积分10
13秒前
小小吴完成签到,获得积分10
16秒前
gyyy完成签到,获得积分10
16秒前
李爱国应助沉静的书南采纳,获得10
17秒前
玺青一生完成签到 ,获得积分10
18秒前
肉卷完成签到 ,获得积分10
19秒前
zzh完成签到 ,获得积分10
20秒前
kenny完成签到,获得积分10
23秒前
irisxiong完成签到,获得积分10
24秒前
25秒前
旺旺完成签到,获得积分10
25秒前
Rice完成签到,获得积分10
25秒前
liang19640908完成签到 ,获得积分0
28秒前
zzc完成签到,获得积分10
28秒前
NexusExplorer应助zzh采纳,获得10
29秒前
LBJ完成签到,获得积分10
32秒前
打打应助欧克采纳,获得10
32秒前
chenmeimei2012完成签到 ,获得积分10
35秒前
ding应助tszjw168采纳,获得10
36秒前
mhy完成签到 ,获得积分10
37秒前
YanKangLee12完成签到 ,获得积分10
38秒前
曾经的嘉熙完成签到 ,获得积分10
39秒前
45秒前
机灵石头完成签到,获得积分10
46秒前
46秒前
daomaihu发布了新的文献求助100
46秒前
烂漫的诗蕊完成签到,获得积分10
51秒前
56秒前
东都哈士奇完成签到,获得积分10
56秒前
56秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
自動車の空力技術 800
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7778444
求助须知:如何正确求助?哪些是违规求助? 9318783
关于积分的说明 20366209
捐赠科研通 7365553
什么是DOI,文献DOI怎么找? 3319210
关于科研通互助平台的介绍 2467170
邀请新用户注册赠送积分活动 2334659