亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

From algorithms to operating room: can large language models master China’s attending anesthesiology exam? a cross-sectional evaluation

子专业 医学 麻醉学 集合(抽象数据类型) 医学教育 机器学习 计算机科学 家庭医学 病理 程序设计语言
作者
Qiyu He,Zhimin Tan,Niu Wang,Dongxu Chen,Xian Zhang,Feng Qin,Jiuhong Yuan
出处
期刊:International Journal of Surgery [Wolters Kluwer]
标识
DOI:10.1097/js9.0000000000003406
摘要

Objective: The performance of large language models (LLMs) in complex clinical reasoning tasks is not well established. This study compares ChatGPT (GPT-3.5, GPT-4) and DeepSeek (DeepSeek-V3, DeepSeek-R1) in the Chinese anesthesiology attending physician examination (CAAPE), aiming to set AI benchmarks in medical assessments and enhance AI-driven medical education. Methods: This cross-sectional study assessed four iterations of two major LLMs on the 2025 CAAPE question bank (5,647 questions). Testing employed diverse querying strategies and languages, with subgroup analyses by subspecialty, knowledge type, and question format. The focus was on LLM performance in clinical and logical reasoning tasks, measuring accuracy, error types, and response times. Results: DeepSeek-R1 (70.6%-73.4%) and GPT-4 (68.6%-70.3%) outperformed DeepSeek-V3 (53.1%-55.5%) and GPT-3.5 (52.2%-55.7%) across all strategies. System role (SR) improved performance, while joint response degraded it. DeepSeek-R1 outperformed GPT-4 in complex subspecialties, reaching peak accuracy (73.4%) under SR combined initial response. GPT models performed better with English than Chinese queries. All models excelled in basic knowledge and Type A1 questions but struggled with clinical scenarios and advanced reasoning. Despite DeepSeek-R1’s stronger performance, its response time was longer. Errors were primarily logical and informational (over 70%), with more than half being high-risk clinical errors. Conclusion: LLMs show promise in complex clinical reasoning but risk critical errors in high-risk settings. While useful for education and decision support, their error potential must be carefully assessed in high-stakes environments.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
科研通AI6.2应助任雨光采纳,获得10
7秒前
12秒前
闭家锁发布了新的文献求助30
15秒前
优秀函完成签到,获得积分10
16秒前
ATREE完成签到,获得积分10
17秒前
华仔应助学术混子采纳,获得10
18秒前
潇洒的大神完成签到,获得积分10
44秒前
1分钟前
米米发布了新的文献求助10
1分钟前
cdercder应助科研通管家采纳,获得10
1分钟前
汉堡包应助科研通管家采纳,获得10
1分钟前
1分钟前
赫连山菡发布了新的文献求助10
1分钟前
柔弱的妙旋完成签到,获得积分10
1分钟前
桐桐应助Pami采纳,获得10
1分钟前
靓丽的山蝶完成签到 ,获得积分10
1分钟前
1分钟前
超帅的幻枫完成签到,获得积分10
1分钟前
2分钟前
2分钟前
Pami发布了新的文献求助10
2分钟前
早123完成签到 ,获得积分10
2分钟前
清新的涵双完成签到 ,获得积分10
2分钟前
2分钟前
切尔茜发布了新的文献求助10
2分钟前
我是老大应助wwwww采纳,获得10
2分钟前
2分钟前
Noob_saibot完成签到,获得积分10
2分钟前
田様应助wwwww采纳,获得30
2分钟前
英姑应助wwwww采纳,获得10
2分钟前
思源应助wwwww采纳,获得10
2分钟前
我是老大应助wwwww采纳,获得10
2分钟前
烟花应助wwwww采纳,获得30
2分钟前
我是老大应助wwwww采纳,获得10
2分钟前
万能图书馆应助wwwww采纳,获得10
2分钟前
orixero应助wwwww采纳,获得30
2分钟前
所所应助wwwww采纳,获得30
2分钟前
CipherSage应助wwwww采纳,获得30
2分钟前
粗心的书竹完成签到,获得积分10
2分钟前
2分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Art Therapy and Career Counseling 600
The Oxford Handbook of Digital Classical Studies 550
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7619130
求助须知:如何正确求助?哪些是违规求助? 9194586
关于积分的说明 19706136
捐赠科研通 7191201
什么是DOI,文献DOI怎么找? 3272388
关于科研通互助平台的介绍 2435003
邀请新用户注册赠送积分活动 2267604