Assessing Large Language Models for Oncology Data Inference From Radiology Reports

推论 医学物理学 计算机科学 医学 数据科学 人工智能
作者
L. Chen,Travis Zack,Arda Demirci,Madhumita Sushil,Brenda Y. Miao,Corynn Kasap,Atul J. Butte,Eric A. Collisson,Julian C. Hong
出处
期刊:JCO clinical cancer informatics [Lippincott Williams & Wilkins]
卷期号:8 (8): e2400126-e2400126 被引量:10
标识
DOI:10.1200/cci.24.00126
摘要

PURPOSE We examined the effectiveness of proprietary and open large language models (LLMs) in detecting disease presence, location, and treatment response in pancreatic cancer from radiology reports. METHODS We analyzed 203 deidentified radiology reports, manually annotated for disease status, location, and indeterminate nodules needing follow-up. Using generative pre-trained transformer (GPT)-4, GPT-3.5-turbo, and open models such as Gemma-7B and Llama3-8B, we employed strategies such as ablation and prompt engineering to boost accuracy. Discrepancies between human and model interpretations were reviewed by a secondary oncologist. RESULTS Among 164 patients with pancreatic tumor, GPT-4 showed the highest accuracy in inferring disease status, achieving a 75.5% correctness (F1-micro). Open models Mistral-7B and Llama3-8B performed comparably, with accuracies of 68.6% and 61.4%, respectively. Mistral-7B excelled in deriving correct inferences from objective findings directly. Most tested models demonstrated proficiency in identifying disease containing anatomic locations from a list of choices, with GPT-4 and Llama3-8B showing near-parity in precision and recall for disease site identification. However, open models struggled with differentiating benign from malignant postsurgical changes, affecting their precision in identifying findings indeterminate for cancer. A secondary review occasionally favored GPT-3.5's interpretations, indicating the variability in human judgment. CONCLUSION LLMs, especially GPT-4, are proficient in deriving oncologic insights from radiology reports. Their performance is enhanced by effective summarization strategies, demonstrating their potential in clinical support and health care analytics. This study also underscores the possibility of zero-shot open model utility in environments where proprietary models are restricted. Finally, by providing a set of annotated radiology reports, this paper presents a valuable data set for further LLM research in oncology.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
fktd发布了新的文献求助10
1秒前
采采完成签到,获得积分10
2秒前
踏实的魔镜完成签到,获得积分10
2秒前
2秒前
小蘑菇应助彩虹追月采纳,获得10
3秒前
大鹏完成签到,获得积分10
3秒前
搬砖一号发布了新的文献求助10
4秒前
youth应助媛小媛啊采纳,获得10
4秒前
meow发布了新的文献求助10
5秒前
蓝天应助朴素海亦采纳,获得10
6秒前
6秒前
yu完成签到,获得积分10
6秒前
今后应助三毛采纳,获得10
7秒前
mr发布了新的文献求助10
8秒前
畅快灵薇发布了新的文献求助10
8秒前
8秒前
淡然太清完成签到,获得积分10
8秒前
青衣完成签到,获得积分10
9秒前
9秒前
11秒前
11秒前
MM完成签到,获得积分10
11秒前
12秒前
xxx发布了新的文献求助10
12秒前
13秒前
尊嘟假嘟发布了新的文献求助100
13秒前
13秒前
cranberry发布了新的文献求助10
14秒前
14秒前
lsblb发布了新的文献求助10
15秒前
zyy完成签到,获得积分20
15秒前
彩虹追月发布了新的文献求助10
16秒前
17秒前
17秒前
纯牛奶发布了新的文献求助10
18秒前
18秒前
阿聪发布了新的文献求助10
18秒前
18秒前
Jasper应助余咋采纳,获得10
20秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Organic Chemistry, 5th Edition 1000
Handbook of Social Psychology and Consumer Behavior 900
Nondestructive Testing Handbook: Vol. 4, Thermal and Infrared Testing (IR), 4th ed 800
日本現代怪異事典 副読本 700
Handbook of Social Identity Research 600
作者名:Kristopher P. Plain,悉尼大学的,目前只能查到其四篇论文,想找到其博士论文 590
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7374496
求助须知:如何正确求助?哪些是违规求助? 8982236
关于积分的说明 19097309
捐赠科研通 7015487
什么是DOI,文献DOI怎么找? 3225685
关于科研通互助平台的介绍 2389020
邀请新用户注册赠送积分活动 2206219