已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

Evaluating large language models as graders of medical short answer questions: a comparative analysis with expert human graders

医学教育 数学教育 心理学 数据科学 计算机科学 医学
作者
Olena Bolgova,Paul Ganguly,Muhammad Faisal Ikram,Volodymyr Mavrych
出处
期刊:Medical Education Online [Taylor & Francis]
卷期号:30 (1)
标识
DOI:10.1080/10872981.2025.2550751
摘要

The assessment of short-answer questions (SAQs) in medical education is resource-intensive, requiring significant expert time. Large Language Models (LLMs) offer potential for automating this process, but their efficacy in specialized medical education assessment remains understudied. To evaluate the capability of five LLMs to grade medical SAQs compared to expert human graders across four distinct medical disciplines. This study analyzed 804 student responses across anatomy, histology, embryology, and physiology. Three faculty members graded all responses. Five LLMs (GPT-4.1, Gemini, Claude, Copilot, DeepSeek) evaluated responses twice: first using their learned representations to generate their own grading criteria (A1), then using expert-provided rubrics (A2). Agreement was measured using Cohen's Kappa and Intraclass Correlation Coefficient (ICC). Expert-expert agreement was substantial across all questions (average Kappa: 0.69, ICC: 0.86), ranging from moderate (SAQ2: 0.57) to almost perfect (SAQ4: 0.87). LLM performance varied dramatically by question type and model. The highest expert-LLM agreement was observed for Claude on SAQ3 (Kappa: 0.61) and DeepSeek on SAQ2 (Kappa: 0.53). Providing expert criteria had inconsistent effects, significantly improving some model-question combinations while decreasing others. No single LLM consistently outperformed others across all domains. LLM strictness in grading unsatisfactory responses varied substantially from experts. LLMs demonstrated domain-specific variations in grading capabilities. The provision of expert criteria did not consistently improve performance. While LLMs show promise for supporting medical education assessment, their implementation requires domain-specific considerations and continued human oversight.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
2秒前
星河入梦来完成签到,获得积分10
4秒前
6秒前
认真不可发布了新的文献求助10
7秒前
小蘑菇应助DSHR采纳,获得10
8秒前
cbb发布了新的文献求助10
12秒前
Tao完成签到 ,获得积分10
13秒前
坦率的语柳完成签到 ,获得积分10
13秒前
灵巧小夏完成签到,获得积分10
14秒前
17秒前
18秒前
吴糖完成签到,获得积分10
19秒前
cbb发布了新的文献求助10
20秒前
22秒前
22秒前
DSHR发布了新的文献求助10
23秒前
高兴的小天鹅完成签到,获得积分10
24秒前
24秒前
26秒前
tiara完成签到 ,获得积分10
27秒前
Bobo发布了新的文献求助10
28秒前
28秒前
28秒前
28秒前
逍遥游233完成签到 ,获得积分10
30秒前
31秒前
西原的橙果完成签到,获得积分10
32秒前
chemcf完成签到,获得积分10
32秒前
33秒前
gt发布了新的文献求助10
34秒前
哈基米曼波完成签到,获得积分10
36秒前
Karel发布了新的文献求助10
36秒前
36秒前
赘婿应助笨笨的灵竹采纳,获得10
37秒前
科研通AI6.4应助cbb采纳,获得10
39秒前
愉快宛凝完成签到,获得积分20
41秒前
关尔匕禾页完成签到,获得积分10
42秒前
aajhajkahna应助羞涩的笑天采纳,获得10
43秒前
nano完成签到 ,获得积分10
43秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Effects of Two Weeks of Red Light Therapy on Choroidal Thickness and Axial Length in Young Adults 700
内視鏡的に摘除しえた十二指腸乳頭部腫瘍の2例 660
The Foundation of Positive Psychology 600
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The Neuroscience of Language 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7676760
求助须知:如何正确求助?哪些是违规求助? 9242710
关于积分的说明 19918682
捐赠科研通 7247041
什么是DOI,文献DOI怎么找? 3286558
关于科研通互助平台的介绍 2444550
邀请新用户注册赠送积分活动 2289570