医学
一致性(知识库)
医学教育
医疗急救
临床决策
急诊科
梅德林
临床决策支持系统
紧急医疗服务
替代医学
医学知识
医学物理学
临床实习
临床判断
决策支持系统
人工智能
作者
İshak Şan,Medine Akkan Öz,Mehmet Yortanlı,Murat Genç,Bensu Bulut,AYŞENUR GÜR,RAMİZ YAZICI,Hüseyin Mutlu,MUSTAFA ÖNDER GÖNEN
标识
DOI:10.55730/1300-0144.6083
摘要
Background/aim: This study evaluated the accuracy rates and response consistency of four different large language models (ChatGPT-4o, Gemini 2.0, Claude 3.5, and DeepSeek R1) in answering questions from the Emergency Medicine Fellowship Examination (YDUS), which was administered for the first time in Türkiye. Materials and methods: In this observational study, 60 multiple-choice questions from the Emergency Medicine YDUS administered on 15 December 2024, were classified as knowledge-based (n = 26), visual content (n = 2), and case-based (n = 32). Each question was presented three times to the four large language models. The models' accuracy rates were evaluated according to overall accuracy, strict accuracy, and ideal accuracy criteria. Response consistency was measured using Fleiss' Kappa test. Results: The ChatGPT-4o model was the most successful in terms of overall accuracy (90.0%), while DeepSeek R1 showed the lowest performance (76.7%). Claude 3.5 (83.3%) and Gemini 2.0 (80.0%) demonstrated moderate success. When analyzed by category, ChatGPT-4o achieved the highest success with 92.3% accuracy in knowledge-based questions and 90.6% in case-based questions. In terms of response consistency, the Claude 3.5 model (Fleiss' Kappa = 0.68) showed the highest consistency, while Gemini 2.0 (Fleiss' Kappa = 0.49) showed the lowest. Inconsistent hallucinations were more frequent in the Gemini 2.0 and DeepSeek R1 models, whereas persistent hallucinations were less common in the ChatGPT-4o and Claude 3.5 models. Conclusion: Large language models can achieve high accuracy rates for knowledge and clinical reasoning questions in emergency medicine but show differences in terms of response consistency and hallucination tendency. While these models have significant potential for use in medical education and as clinical decision support systems (CDSS), they need further development to provide reliable, up-to-date, and accurate information.
科研通智能强力驱动
Strongly Powered by AbleSci AI