医学诊断
计算机科学
人工智能
软件部署
基于模型的推理
认知心理学
边疆
过程(计算)
自然语言处理
言语推理
心理学
差速器(机械装置)
基于案例的推理
决策模型
风险分析(工程)
因果推理
证据推理法
推理系统
作者
Arya S. Rao,Kaiz P. Esmail,Richard S. Lee,Sharon Jiang,Bianca Arraiza Carlo,Jasleen Gill,Praneet Khanna,Ezra Kalmowitz,Basile Montagnese,Kimia Heydari,Qiao Jiao,Ethan Bott,Dan Nguyen,Grace Wang,Michael Hood,Adam Landman,Marc D. Succi
出处
期刊:JAMA network open
[American Medical Association]
日期:2026-04-13
卷期号:9 (4): e264003-e264003
标识
DOI:10.1001/jamanetworkopen.2026.4003
摘要
In this cross-sectional study of 21 LLMs, frontier LLMs achieved high accuracy on final diagnoses but performed poorly in generating differential diagnoses and navigating uncertainty relative to other reasoning stages. The PrIME-LLM framework provided greater separation than raw accuracy, revealing critical reasoning gaps obscured by traditional benchmarks. Thus, despite version-based improvements and advantages in reasoning-optimized models, off-the-shelf LLMs have not yet achieved the intelligence required for safe deployment and remain limited in demonstrating advanced clinical reasoning.
科研通智能强力驱动
Strongly Powered by AbleSci AI