Large Language Models in Medical Education: Comparing ChatGPT- to Human-Generated Exam Questions

医学教育 心理学 梅德林 高等教育 计算机科学 医学 政治学 法学 生物化学 生物
作者
Matthias Carl Laupichler,Johanna Flora Rother,Ilona C Grunwald Kadow,Seifollah Ahmadi,Tobias Raupach
出处
期刊:Academic Medicine [Lippincott Williams & Wilkins]
卷期号:99 (5): 508-512 被引量:84
标识
DOI:10.1097/acm.0000000000005626
摘要

PROBLEM: Creating medical exam questions is time consuming, but well-written questions can be used for test-enhanced learning, which has been shown to have a positive effect on student learning. The automated generation of high-quality questions using large language models (LLMs), such as ChatGPT, would therefore be desirable. However, there are no current studies that compare students' performance on LLM-generated questions to questions developed by humans. APPROACH: The authors compared student performance on questions generated by ChatGPT (LLM questions) with questions created by medical educators (human questions). Two sets of 25 multiple-choice questions (MCQs) were created, each with 5 answer options, 1 of which was correct. The first set of questions was written by an experienced medical educator, and the second set was created by ChatGPT 3.5 after the authors identified learning objectives and extracted some specifications from the human questions. Students answered all questions in random order in a formative paper-and-pencil test that was offered leading up to the final summative neurophysiology exam (summer 2023). For each question, students also indicated whether they thought it had been written by a human or ChatGPT. OUTCOMES: The final data set consisted of 161 participants and 46 MCQs (25 human and 21 LLM questions). There was no statistically significant difference in item difficulty between the 2 question sets, but discriminatory power was statistically significantly higher in human than LLM questions (mean = .36, standard deviation [SD] = .09 vs mean = .24, SD = .14; P = .001). On average, students identified 57% of question sources (human or LLM) correctly. NEXT STEPS: Future research should replicate the study procedure in other contexts (e.g., other medical subjects, semesters, countries, and languages). In addition, the question of whether LLMs are suitable for generating different question types, such as key feature questions, should be investigated.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
大气迎天完成签到,获得积分10
1秒前
真找不到完成签到,获得积分10
1秒前
王金娥完成签到,获得积分10
1秒前
Juvenilesy应助ale采纳,获得10
2秒前
hongtaoli2024完成签到 ,获得积分10
2秒前
王京华完成签到,获得积分10
2秒前
sophia完成签到 ,获得积分10
2秒前
奕柯完成签到,获得积分10
2秒前
杰_骜不驯完成签到,获得积分10
3秒前
julia发布了新的文献求助10
3秒前
湖里完成签到,获得积分10
3秒前
科研通AI6.4应助大只00采纳,获得10
3秒前
YY完成签到,获得积分10
4秒前
4秒前
木木完成签到,获得积分10
4秒前
邵初蓝完成签到,获得积分10
4秒前
5秒前
都要多喝水完成签到,获得积分10
5秒前
啵啵完成签到,获得积分0
5秒前
ZhaoQingnai完成签到,获得积分10
6秒前
瓦力文完成签到,获得积分10
6秒前
无花果应助小殷采纳,获得10
6秒前
英勇冰蓝完成签到,获得积分10
6秒前
MCs完成签到,获得积分10
7秒前
拉长的芷烟完成签到 ,获得积分10
7秒前
LLLLL完成签到,获得积分10
7秒前
丁小二完成签到 ,获得积分10
8秒前
8秒前
Kevin完成签到,获得积分10
8秒前
疯子发布了新的文献求助10
9秒前
hhm完成签到,获得积分10
9秒前
9秒前
JA完成签到,获得积分10
10秒前
Georgecat完成签到,获得积分10
10秒前
无心的星月完成签到 ,获得积分10
10秒前
Wu发布了新的文献求助10
10秒前
冷艳的班完成签到,获得积分10
11秒前
思源应助颜林林采纳,获得10
11秒前
铁锅炖大鹅完成签到,获得积分10
11秒前
姜小时完成签到,获得积分10
11秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
化工安全与环保 1000
Autoparametric Resonance in Mechanical Systems 1000
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 800
Effects of Two Weeks of Red Light Therapy on Choroidal Thickness and Axial Length in Young Adults 700
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7656698
求助须知:如何正确求助?哪些是违规求助? 9227402
关于积分的说明 19829554
捐赠科研通 7223232
什么是DOI,文献DOI怎么找? 3280353
关于科研通互助平台的介绍 2440621
邀请新用户注册赠送积分活动 2280188