清晨好,您是今天最早来到科研通的研友!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您科研之路漫漫前行!

Lung Cancer Staging Using Chest CT and FDG PET/CT Free-Text Reports: Comparison Among Three ChatGPT Large Language Models and Six Human Readers of Varying Experience

医学 肺癌 放射科 癌症 正电子发射断层摄影术 医学物理学 病理 内科学
作者
Jong Eun Lee,Ki Seong Park,Yun‐Hyeon Kim,Ho-Chun Song,Byunggeon Park,Yeon Joo Jeong
出处
期刊:American Journal of Roentgenology [American Roentgen Ray Society]
卷期号:223 (6): e2431696-e2431696 被引量:26
标识
DOI:10.2214/ajr.24.31696
摘要

BACKGROUND. Although radiology reports are commonly used for lung cancer staging, this task can be challenging given radiologists' variable reporting styles as well as reports' potentially ambiguous and/or incomplete staging-related information. OBJECTIVE. The purpose of this study was to compare the performance of ChatGPT large language models (LLMs) and human readers of varying experience in lung cancer staging using chest CT and FDG PET/CT free-text reports. METHODS. This retrospective study included 700 patients (mean age, 73.8 ± 29.5 [SD] years; 509 men, 191 women) from four institutions in Korea who underwent chest CT or FDG PET/CT for non-small cell lung cancer initial staging from January 2020 to December 2023. Examinations' reports used a free-text format, written exclusively in English or in mixed English and Korean. Two thoracic radiologists in consensus determined the overall stage group (IA, IB, IIA, IIB, IIIA, IIIB, IIIC, IVA, or IVB) for each report using the 8th-edition AJCC Cancer Staging Manual to establish the reference standard. Three ChatGPT models (GPT-4o, GPT-4, GPT-3.5) determined an overall stage group for each report using a script-based application programming interface, zero-shot learning, and a prompt incorporating a staging system summary. The code for this web application was made publicly available through a GitHub repository (https://github.com/elmidion/GPT_Information_Extractor). Six human readers (two fellowship-trained radiologists with less experience than the radiologists who determined the reference standard, two fellows, and two residents) also independently determined overall stage groups. GPT-4o's overall accuracy for determining the correct stage among the nine groups was compared with that of the other LLMs and human readers using McNemar tests. RESULTS. GPT-4o had an overall staging accuracy of 74.1%, significantly better than the accuracy of GPT-4 (70.1%, p = .02), GPT-3.5 (57.4%, p < .001), and resident 2 (65.7%, p < .001); significantly worse than the accuracy of fellowship-trained radiologist 1 (82.3%, p < .001) and fellowship-trained radiologist 2 (85.4%, p < .001); and not significantly different from the accuracy of fellow 1 (77.7%, p = .09), fellow 2 (75.6%, p = .53), and resident 1 (72.3%, p = .42). CONCLUSION. The best-performing model, GPT-4o, showed no significant difference in staging accuracy versus fellows but showed significantly worse performance versus fellowship-trained radiologists. The findings do not support use of LLMs for lung cancer staging in place of expert health care professionals. CLINICAL IMPACT. The findings indicate the importance of domain expertise for performing complex specialized tasks such as cancer staging.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
枫威完成签到 ,获得积分10
1秒前
1秒前
王亚楠完成签到 ,获得积分10
4秒前
whitepiece完成签到,获得积分0
10秒前
11秒前
cheng完成签到,获得积分10
17秒前
油麦集团大小姐完成签到 ,获得积分10
19秒前
Kiry完成签到 ,获得积分10
21秒前
22秒前
zy完成签到,获得积分10
26秒前
足下慵才完成签到,获得积分10
43秒前
辛勤爆米花给一朵懒云的求助进行了留言
43秒前
kkk完成签到 ,获得积分10
45秒前
Harlotte完成签到 ,获得积分10
54秒前
波西米亚完成签到,获得积分10
55秒前
桐桐应助NaiZeMu采纳,获得10
1分钟前
夏日随笔完成签到 ,获得积分10
1分钟前
林好人完成签到 ,获得积分10
1分钟前
1分钟前
felicity完成签到 ,获得积分10
1分钟前
干净的绿真完成签到 ,获得积分10
1分钟前
赵赶超发布了新的文献求助10
1分钟前
felicia12138完成签到 ,获得积分10
1分钟前
1分钟前
DrHHB完成签到 ,获得积分10
1分钟前
青水完成签到 ,获得积分10
1分钟前
1分钟前
1分钟前
cquank完成签到,获得积分10
1分钟前
小鱼崽完成签到 ,获得积分10
1分钟前
fizzy完成签到 ,获得积分10
1分钟前
1分钟前
leapper完成签到 ,获得积分10
1分钟前
1分钟前
养花低手完成签到 ,获得积分0
2分钟前
研友_ZzrWKZ完成签到 ,获得积分10
2分钟前
想要吉姆尼完成签到,获得积分10
2分钟前
XRH完成签到,获得积分10
2分钟前
2分钟前
鑫鑫完成签到,获得积分10
2分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Les chinois de jakarta: temples et vie collective 1000
Autoparametric Resonance in Mechanical Systems 1000
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 800
Social Psychology 600
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7646079
求助须知:如何正确求助?哪些是违规求助? 9218401
关于积分的说明 19778026
捐赠科研通 7210505
什么是DOI,文献DOI怎么找? 3276958
关于科研通互助平台的介绍 2438619
邀请新用户注册赠送积分活动 2275023