计算机科学
人工智能
基线(sea)
金标准(测试)
专家系统
语言模型
自然语言处理
医学教育
梅德林
基于模型的推理
机器学习
言语推理
临床决策支持系统
病历
心理学
医学
基于案例的推理
临床决策
数据科学
急诊科
决策支持系统
医疗急救
科学推理
作者
Peter G. Brodeur,Thomas A Buckley,Zahir Kanjee,Ethan Goh,Evelyn Ling,Priyank Jain,Stephanie Cabral,Raja-Elie Abdulnour,Adrian D. Haimovich,Jason A. Freed,Andrew Olson,Daniel J Morgan,Jason Hom,Robert Gallo,Liam G McCoy,Haadi Mombini,Christopher Lucas,M. Fotoohi,Matthew Gwiazdon,Daniele Restifo
出处
期刊:Science
[American Association for the Advancement of Science]
日期:2026-04-30
卷期号:392 (6797): 524-527
被引量:28
标识
DOI:10.1126/science.adz4433
摘要
More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. We then report a real-world study comparing human expert and artificial intelligence (AI) second opinions in randomly selected patients in the emergency room of a major tertiary academic medical center. In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials.
科研通智能强力驱动
Strongly Powered by AbleSci AI