Structured Report Generation for Breast Cancer Imaging Based on Large Language Modeling: A Comparative Analysis of GPT-4 and DeepSeek

麦克内马尔试验 医学 放射科 乳腺癌 乳腺摄影术 病变 一致性 乳房成像 癌症 医学物理学 内科学 病理 统计 数学
作者
Kun Chen,Xuefeng Hou,Xiaofeng Li,Wengui Xu,Heqing Yi
出处
期刊:Academic Radiology [Elsevier BV]
卷期号:32 (10): 5693-5702 被引量:10
标识
DOI:10.1016/j.acra.2025.07.046
摘要

RATIONALE AND OBJECTIVES: The purpose of this study is to compare the performance of GPT-4 and DeepSeek large language models in generating structured breast cancer multimodality imaging integrated reports from free-text radiology reports including mammography, ultrasound, MRI, and PET/CT. MATERIALS AND METHODS: A retrospective analysis was conducted on 1358 free-text reports from 501 breast cancer patients across two institutions. The study design involved synthesizing multimodal imaging data into structured reports with three components: primary lesion characteristics, metastatic lesions, and TNM staging. Input prompts were standardized for both models, with GPT-4 using predesigned instructions and DeepSeek requiring manual input. Reports were evaluated based on physician satisfaction using a Likert scale, descriptive accuracy including lesion localization, size, SUV, and metastasis assessment, and TNM staging correctness according to NCCN guidelines. Statistical analysis included McNemar tests for binary outcomes and correlation analysis for multiclass comparisons with a significance threshold of P < .05. RESULTS: Physician satisfaction scores showed strong correlation between models with r-values of 0.665 and 0.558 and P-values below .001. Both models demonstrated high accuracy in data extraction and integration. The mean accuracy for primary lesion features was 91.7% for GPT-4% and 92.1% for DeepSeek, while feature synthesis accuracy was 93.4% for GPT4 and 93.9% for DeepSeek. Metastatic lesion identification showed comparable overall accuracy at 93.5% for GPT4 and 94.4% for DeepSeek. GPT-4 performed better in pleural lesion detection with 94.9% accuracy compared to 79.5% for DeepSeek, whereas DeepSeek achieved higher accuracy in mesenteric metastasis identification at 87.5% vs 43.8% for GPT4. TNM staging accuracy exceeded 92% for T-stage and 94% for M-stage, with N-stage accuracy improving beyond 90% when supplemented with physical exam data. CONCLUSION: Both GPT-4 and DeepSeek effectively generate structured breast cancer imaging reports with high accuracy in data mining, integration, and TNM staging. Integrating these models into clinical practice is expected to enhance report standardization and physician productivity.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
心想事成完成签到,获得积分10
刚刚
叶祥发布了新的文献求助10
刚刚
北辰发布了新的文献求助10
1秒前
yanni应助z张采纳,获得10
1秒前
1秒前
东方不败完成签到 ,获得积分20
1秒前
1秒前
Criminology34应助涵涵可以采纳,获得30
2秒前
2秒前
充电宝应助六元一斤虾采纳,获得10
3秒前
xuan发布了新的文献求助10
3秒前
wyy关闭了wyy文献求助
4秒前
4秒前
sagitar应助gyusbjshaxb采纳,获得40
4秒前
Jasper应助李木槿采纳,获得10
5秒前
桐桐应助冻梨采纳,获得10
6秒前
爆米花应助六月采纳,获得10
6秒前
7秒前
李wf完成签到,获得积分10
7秒前
Dorjee发布了新的文献求助10
7秒前
8秒前
囧囧给牧青的求助进行了留言
9秒前
9秒前
9秒前
11秒前
monica发布了新的文献求助10
12秒前
leeyc发布了新的文献求助10
14秒前
14秒前
Cici发布了新的文献求助30
16秒前
欢呼醉薇发布了新的文献求助10
16秒前
18秒前
19秒前
Mei完成签到,获得积分10
20秒前
李木槿发布了新的文献求助10
20秒前
drinkliu完成签到,获得积分10
21秒前
看满天星河完成签到 ,获得积分10
23秒前
HLQF完成签到,获得积分10
23秒前
25秒前
双shuang完成签到,获得积分10
26秒前
显摆小狂人完成签到,获得积分10
26秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7774238
求助须知:如何正确求助?哪些是违规求助? 9316144
关于积分的说明 20349706
捐赠科研通 7359972
什么是DOI,文献DOI怎么找? 3317404
关于科研通互助平台的介绍 2465884
邀请新用户注册赠送积分活动 2332640