Solving Complex Pediatric Surgical Case Studies: A Comparative Analysis of Copilot, ChatGPT-4 and Experienced Pediatric Surgeons’ Performance

医学 普通外科 儿科外科医生 小儿外科 外科
作者
Richard Gnatzy,Martin Lacher,Michael Berger,Michael Boettcher,Oliver Johannes Deffaa,Joachim Kübler,Omid Madadi‐Sanjani,Illya Martynov,Steffi Mayer,Mikko P. Pakarinen,Richard Wagner,Tomas Wester,Augusto Zani,Ophelia Aubert
出处
期刊:European Journal of Pediatric Surgery [Thieme Medical Publishers (Germany)]
标识
DOI:10.1055/a-2551-2131
摘要

The emergence of large language models (LLMs) has led to notable advancements across multiple sectors, including medicine. Yet, their effect in pediatric surgery remains largely unexplored. This study aims to assess the ability of the AI models ChatGPT-4 and Microsoft Copilot to propose diagnostic procedures, primary and differential diagnoses, as well as answer clinical questions using complex clinical case vignettes of classic pediatric surgical diseases. We conducted the study in April 2024. We evaluated the performance of LLMs using 13 complex clinical case vignettes of pediatric surgical diseases and compared responses to a human cohort of experienced pediatric surgeons. Additionally, pediatric surgeons rated the diagnostic recommendations of LLMs for completeness and accuracy. To determine differences in performance we performed statistical analyses. ChatGPT-4 achieved a higher test score (52.1%) compared to Copilot (47.9%), but less than pediatric surgeons (68.8%). Overall differences in performance between ChatGPT-4, Copilot, and pediatric surgeons were found to be statistically significant (p <0.01). ChatGPT-4 demonstrated a superior performance in generating differential diagnoses compared to Copilot (p<0.05). No statistically significant differences were found between the AI models regarding suggestions for diagnostics and primary diagnosis. Overall, recommendations of LLMs were rated as average by pediatric surgeons. This study reveals significant limitations in the performance of AI models in pediatric surgery. Although LLMs exhibit potential across various areas, their reliability and accuracy in handling clinical decision-making tasks is limited. Further research is needed to improve AI capabilities and establish its usefulness in the clinical setting.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
panxixiang发布了新的文献求助10
刚刚
刚刚
端一眼完成签到,获得积分20
刚刚
11完成签到,获得积分10
刚刚
刚刚
wenwen完成签到 ,获得积分10
刚刚
熊先生完成签到,获得积分10
1秒前
1秒前
1秒前
领导范儿的应助被诚心花生采纳,获得10
2秒前
ACE完成签到,获得积分10
2秒前
Fanss发布了新的文献求助10
2秒前
DW的应助被小v1212采纳,获得10
3秒前
3秒前
HAO发布了新的文献求助10
4秒前
XIEYIHAN发布了新的文献求助10
4秒前
多洛塔完成签到 ,获得积分10
4秒前
xzj7789210完成签到,获得积分10
4秒前
cxh完成签到,获得积分10
4秒前
orixero的应助被笑点低的小笼包采纳,获得10
4秒前
4秒前
4秒前
fangyuan完成签到,获得积分10
5秒前
小灰灰完成签到 ,获得积分10
5秒前
5秒前
5秒前
努力学好完成签到,获得积分10
6秒前
6秒前
6秒前
爆米花的应助被wjx采纳,获得10
7秒前
7秒前
7秒前
7秒前
Yumion完成签到,获得积分10
8秒前
lulu发布了新的文献求助10
8秒前
panxixiang完成签到,获得积分20
8秒前
8秒前
卖艺的读书人完成签到 ,获得积分10
8秒前
Key发布了新的文献求助10
8秒前
青天鸟1989完成签到,获得积分10
9秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
The Art of Interactive Teaching 600
Computational Chemical Reaction Engineering: Modeling, Simulation, and Design with MATLAB 600
Organizational Behavior 510
Management and the Arts 510
CLSI C56QG Examples of Hemolyzed, Icteric, and Lipemic/Turbid Samples Quick Guide 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 内科学 物理 有机化学 化学工程 生物化学 复合材料 光电子学 细胞生物学 心理学 量子力学 催化作用 物理化学 电极
热门帖子
关注 科研通微信公众号,转发送积分 7799781
求助须知:如何正确求助?哪些是违规求助? 9334816
关于积分的说明 20469553
捐赠科研通 7390994
什么是DOI,文献DOI怎么找? 3326207
关于科研通互助平台的介绍 2473161
邀请新用户注册赠送积分活动 2343866