GPT-4 vs. Radiologists: who advances mediastinal tumor classification better across report quality levels? A cohort study

医学 放射性武器 放射科 诊断准确性 核医学
作者
Ru Wen,Xiaoming Li,Kang Chen,Mo Sun,Chunxia Zhu,Peng Xu,Fengxi Chen,Chueng‐Ryong Ji,P Groot MI,Xuefeng Li,Xiaojuan Deng,Quan Yang,Weixiang Song,Yajun Shang,Huang Sheng,Mingyang Zhou,Jian Wang,Chaoyang Zhou,Wei Chen,Liu Chen
出处
期刊:International Journal of Surgery [Wolters Kluwer]
标识
DOI:10.1097/js9.0000000000003127
摘要

Background: Accurate mediastinal tumor classification is crucial for treatment planning, but diagnostic performance varies with radiologists’ experience and report quality. Purpose: To evaluate GPT-4’s diagnostic accuracy in classifying mediastinal tumors from radiological reports compared to radiologists of different experience levels using radiological reports of varying quality. Materials and Methods: We conducted a retrospective study of 1,494 patients from five tertiary hospitals with mediastinal tumors diagnosed via chest CT and pathology. Radiological reports were categorized into low-, medium-, and high-quality based on predefined criteria assessed by experienced radiologists. Six radiologists (two residents, two attending radiologists, and two associate senior radiologists) and GPT-4 evaluated the chest CT reports. Diagnostic performance was analyzed overall, by report quality, and by tumor type using Wald χ 2 tests and 95% CIs calculated via the Wilson method. Results: GPT-4 achieved an overall diagnostic accuracy of 73.3% (95% CI: 71.0–75.5), comparable to associate senior radiologists (74.3%, 95% CI: 72.0–76.5; p >0.05). For low-quality reports, GPT-4 outperformed associate senior radiologists (60.8% vs. 51.1%, p <0.001). In high-quality reports, GPT-4 was comparable to attending radiologists (80.6% vs.79.4%, p >0.05). Diagnostic performance varied by tumor type: GPT-4 was comparable to radiology residents for neurogenic tumors (44.9% vs. 50.3%, p >0.05), similar to associate senior radiologists for teratomas (68.1% vs. 65.9%, p >0.05), and superior in diagnosing lymphoma (75.4% vs. 60.4%, p <0.001). Conclusion: GPT-4 demonstrated interpretation accuracy comparable to Associate Senior Radiologists, excelling in low-quality reports and outperforming them in diagnosing lymphoma. These findings underscore GPT-4’s potential to enhance diagnostic performance in challenging diagnostic scenarios.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
ShaeRay完成签到,获得积分10
刚刚
刚刚
毛毛虫完成签到,获得积分10
刚刚
1秒前
美满的酸奶完成签到,获得积分10
1秒前
1秒前
赘婿应助oioioi采纳,获得15
2秒前
djfndnn完成签到 ,获得积分10
3秒前
小陈完成签到,获得积分10
4秒前
梓渝完成签到 ,获得积分10
4秒前
赵春晖发布了新的文献求助10
4秒前
5秒前
mingxing818完成签到,获得积分10
5秒前
lemon发布了新的文献求助10
5秒前
yfzhang发布了新的文献求助10
5秒前
依旧完成签到,获得积分10
6秒前
cua发布了新的文献求助10
6秒前
6秒前
顺利的妖妖完成签到 ,获得积分10
6秒前
乃惜发布了新的文献求助10
6秒前
111完成签到,获得积分10
7秒前
8秒前
8秒前
浮华乱世应助windtalker采纳,获得100
8秒前
好运发布了新的文献求助10
9秒前
9秒前
菠萝冰完成签到,获得积分10
9秒前
10秒前
冷静的尔白完成签到,获得积分10
10秒前
10秒前
萨芬完成签到,获得积分10
11秒前
析木发布了新的文献求助10
12秒前
12秒前
zhou应助阿萨德采纳,获得10
13秒前
13秒前
不必要再讨论适合与否完成签到,获得积分0
13秒前
吃皮发布了新的文献求助10
13秒前
13秒前
阳光发布了新的文献求助10
14秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
核安全综合知识2024版 500
Photothermal Science and Techniques 500
Digital Displacement Hydrostatic Transmission for Rotorcraft and Distributed Propulsion 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7711692
求助须知:如何正确求助?哪些是违规求助? 9267972
关于积分的说明 20069426
捐赠科研通 7288365
什么是DOI,文献DOI怎么找? 3297344
关于科研通互助平台的介绍 2451829
邀请新用户注册赠送积分活动 2304349