The effect of data diversity on the performance of deep learning models for predicting early gastric cancer under endoscopy

作者
Conghui Shi,Jia Li,Lianlian Wu
标识
DOI:10.55976/jdh.1202214319-24
摘要

Aims To explore the effect of training set diversity on the performance of deep learning models for predicting early gastric cancer (EGC) under endoscopy. Methods Images of EGC and non-cancerous lesions under narrow-band imaging (ME-NBI) and magnifying blue laser imaging (ME-BLI) were retrospectively collected. Training set 1 was composed of 150 non-cancerous and 309 EGC ME-NBI images, training set 2 was composed of 1505 non-cancerous and 309 EGC ME-BLI images, and training set 3 was the combination of training set 1 and 2. Test set 1 was composed of 376 non-cancerous and 1052 EGC ME-NBI images, test set 2 consisted of 529 non-cancerous and 71 EGC ME-BLI images, and test set 3 was the combination of test set 1 and test set 2. Three deep learning models were constructed, which were respectively CNN 1, CNN 2, and CNN 3 (CNN 1, CNN 2 and CNN 3 were independently trained using training set 1, training set 2 and training set 3 respectively), and their performance on each test set was respectively evaluated. One hundred and thirty-eight ME-NBI videos and 17 ME-BLI videos were further collected to evaluate and compare the performance of each model in real-time. Results On the whole, the performance of CNN 3 was the best. The accuracy (Acc), sensitivity (Sn), specificity (Sp), and area under the curve (AUC) of test set 1 in CNN 3 were 87.89% (1255/1428), 90.96% (342/376), 86.79% (913/1052), and 94.60% respectively. The Acc, Sn, Sp, and AUC of test set 2 in CNN 3 were 95% (570/600), 97.92% (518/529), 73.24% (52/71), and 90.93% respectively. The Acc, Sn, Sp, and AUC of test set 3 in CNN 3 were 89.99% (1825/2028), 95.03% (860/905), 85.93% (965/1123), 94.89% respectively. The performance of CNN 3 was also the best in videos test set. The Acc, Sn, and Sp of videos test set in CNN 3 were 91.03% (142/156), 90.58% (125/138), and 94.44% (17/18) respectively. Conclusions The deep learning model with the most diverse training data has the best diagnostic effect.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
吉瑞完成签到,获得积分20
刚刚
放手一搏完成签到,获得积分10
1秒前
kkk321完成签到,获得积分10
1秒前
彭于晏应助LLL采纳,获得10
1秒前
caowen完成签到 ,获得积分10
1秒前
哄哄发布了新的文献求助10
2秒前
小星完成签到,获得积分10
2秒前
2秒前
天天快乐应助HanYang采纳,获得10
3秒前
3秒前
Orange应助夕瑶采纳,获得10
3秒前
搜集达人应助小超人采纳,获得10
4秒前
英俊的铭应助陈栩采纳,获得10
4秒前
4秒前
4秒前
4秒前
muyu完成签到,获得积分10
5秒前
称心的无色完成签到,获得积分20
5秒前
5秒前
时不言完成签到 ,获得积分10
6秒前
安详的冰凡完成签到 ,获得积分10
7秒前
科研通AI2S应助lh采纳,获得50
7秒前
7秒前
1234完成签到,获得积分10
7秒前
紫色奶萨发布了新的文献求助10
7秒前
科研通AI6.4应助鱼蛋采纳,获得10
7秒前
8秒前
天真雅寒完成签到,获得积分10
8秒前
彭于晏应助拼搏半梦采纳,获得10
8秒前
水之形完成签到,获得积分10
9秒前
王思棋应助文件撤销了驳回
9秒前
Be-a rogue发布了新的文献求助10
9秒前
Baylin发布了新的文献求助10
10秒前
10秒前
11秒前
11秒前
11秒前
11秒前
12秒前
Amber完成签到,获得积分10
12秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Reducing Compassion Fatigue, Secondary Traumatic Stress and Burnout 600
Comparative Elite Sport Development Systems, Structures and Public Policy 600
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Auslegungsgeschichte 500
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 500
What is the Future of Psychotherapy in Digital Age? Technology, AI Bots, and Psychotherapy after Covid 444
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7636632
求助须知:如何正确求助?哪些是违规求助? 9210381
关于积分的说明 19755738
捐赠科研通 7204205
什么是DOI,文献DOI怎么找? 3275510
关于科研通互助平台的介绍 2437215
邀请新用户注册赠送积分活动 2272596