亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

Large-Scale Evaluation of Five Large Language Models in Anesthesia Decision-Making for Hip Fracture Surgery

医学 围手术期 髋部骨折 逻辑回归 梅德林 外科 循证医学 围手术期医学 周围神经 英语 局部麻醉 情感(语言学) 风险评估 血管外科 重症监护医学 麻醉 体格检查 患者安全 神经阻滞 保守管理 并发症 临床实习 物理疗法 病史 麻醉学 随机对照试验 神经轴阻滞
作者
Robert Chen,Andrew Warburton,Ron Do,Darwin D. Chen,Daniel Katz,Garrett W. Burnett
出处
期刊:Anesthesia & Analgesia [Lippincott Williams & Wilkins]
标识
DOI:10.1213/ane.0000000000008148
摘要

BACKGROUND: Large language models (LLMs) show promise for perioperative decision support, but persistent issues, including hallucinations, miscalibration, and biases, indicate they require rigorous evaluation before clinical use. As LLM adoption increases in perioperative settings, systematic evaluation is needed to determine how patient and surgical factors affect performance. METHODS: We evaluated five general-purpose LLMs (DeepSeek 3.2, Gemini 2.5 Flash, GPT-5, GPT-5 mini, GPT-5 nano) using 216 standardized hip fracture surgery vignettes crossing six surgery types, two sexes, and 18 patient variables. We generated 50 samples per combination for 54,000 total responses and collected both structured recommendations and free-text justifications. We used logistic regression to estimate effects on three primary outcomes: anesthesia type, peripheral nerve block placement, and arterial line placement. In a limited sensitivity analysis, we evaluated two clinical LLMs (OpenEvidence, Doximity GPT) with 36 responses each. RESULTS: All models favored neuraxial over general anesthesia (76.1%-88.6% of responses), and all but DeepSeek 3.2 appropriately adjusted recommendations for relevant medical contraindications. All models except GPT-5 nano recommended preoperative peripheral nerve blocks (92.6%-99.3%) and were appropriately conservative regarding arterial line placement. However, free-text justifications frequently cited neuraxial benefits unsupported by recent randomized trials, and most models issued strong neuraxial recommendations despite a lack of clinical justification. We identified limited sociodemographic biases, with only one significant and clinically meaningful effect across 150 comparisons. Clinical LLMs provided similar recommendations to general-purpose models. CONCLUSIONS: While LLMs provided generally reasonable recommendations, systematic preferences diverging from contemporary evidence suggest uncritical use could shift practice patterns without improving patient outcomes.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
李木头完成签到,获得积分10
2秒前
香蕉觅云应助饱满如风采纳,获得10
20秒前
听听完成签到,获得积分10
22秒前
无奈的琦完成签到,获得积分10
25秒前
26秒前
26秒前
复杂鸵鸟完成签到,获得积分10
27秒前
蓝朱发布了新的文献求助30
30秒前
饱满如风发布了新的文献求助10
32秒前
irene完成签到,获得积分10
35秒前
李爱国应助陈运气采纳,获得10
40秒前
睡不醒发布了新的文献求助10
40秒前
52秒前
56秒前
陈运气发布了新的文献求助10
57秒前
清脆的惜萍完成签到,获得积分10
1分钟前
1分钟前
神勇千秋完成签到,获得积分10
1分钟前
1分钟前
平常以云完成签到 ,获得积分10
1分钟前
1分钟前
小蘑菇应助KKSun采纳,获得10
1分钟前
1分钟前
Kao应助科研通管家采纳,获得10
1分钟前
Kao应助科研通管家采纳,获得10
1分钟前
Kao应助科研通管家采纳,获得10
1分钟前
谦让的鹤轩完成签到,获得积分10
1分钟前
活力傲柏完成签到,获得积分10
2分钟前
Owen应助昏睡的金毛采纳,获得10
2分钟前
2分钟前
2分钟前
幸福海之完成签到,获得积分10
2分钟前
如意紫翠完成签到,获得积分10
2分钟前
2分钟前
傲娇书双发布了新的文献求助10
2分钟前
3分钟前
单身的涫完成签到,获得积分10
3分钟前
luo发布了新的文献求助10
3分钟前
顾矜应助luo采纳,获得10
3分钟前
3分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 500
What is the Future of Psychotherapy in Digital Age? Technology, AI Bots, and Psychotherapy after Covid 444
Management and the Arts 310
Teaching Social and Emotional Learning in Physical Education 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7633712
求助须知:如何正确求助?哪些是违规求助? 9207872
关于积分的说明 19748106
捐赠科研通 7202236
什么是DOI,文献DOI怎么找? 3274994
关于科研通互助平台的介绍 2436914
邀请新用户注册赠送积分活动 2271826