Artificial Intelligence as a Discriminator of Competence in Urological Training: Are We There?

鉴别器 医学 能力(人力资源) 医学教育 医学物理学 管理 探测器 电气工程 工程类 经济
作者
Naji J. Touma,Ruchit V. Patel,Thomas A. A. Skinner,Michael Leveridge
出处
期刊:The Journal of Urology [Lippincott Williams & Wilkins]
卷期号:213 (4): 504-511 被引量:7
标识
DOI:10.1097/ju.0000000000004357
摘要

PURPOSE: Assessments in medical education play a central role in evaluating trainees' progress and eventual competence. Generative artificial intelligence is finding an increasing role in clinical care and medical education. The objective of this study was to evaluate the ability of the large language model ChatGPT to generate examination questions that are discriminating in the evaluation of graduating urology residents. MATERIALS AND METHODS: Graduating urology residents representing all Canadian training programs gather yearly for a mock examination that simulates their upcoming board certification examination. The examination consists of a written multiple-choice question (MCQ) examination and an oral objective structured clinical examination. In 2023, ChatGPT Version 4 was used to generate 20 MCQs that were added to the written component. ChatGPT was asked to use Campbell-Walsh Urology, AUA, and Canadian Urological Association guidelines as resources. Psychometric analysis of the ChatGPT MCQs was conducted. The MCQs were also researched by 3 faculty for face validity and to ascertain whether they came from a valid source. RESULTS: The mean score of the 35 examination takers on the ChatGPT MCQs was 60.7% vs 61.1% for the overall examination. Twenty-five of ChatGPT MCQs showed a discrimination index > 0.3, the threshold for questions that properly discriminate between high and low examination performers. Twenty-five percent of ChatGPT MCQs showed a point biserial > 0.2, which is considered a high correlation with overall performance on the examination. The assessment by faculty found that ChatGPT MCQs often provided incomplete information in the stem, provided multiple potentially correct answers, and were sometimes not rooted in the literature. Thirty-five percent of the MCQs generated by ChatGPT provided wrong answers to stems. CONCLUSIONS: Despite what seems to be similar performance on ChatGPT MCQs and the overall examination, ChatGPT MCQs tend not to be highly discriminating. Poorly phrased questions with potential for artificial intelligence hallucinations are ever present. Careful vetting for quality of ChatGPT questions should be undertaken before their use on assessments in urology training examinations.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
桐桐的应助被俏皮眼睛采纳,获得10
刚刚
DXB发布了新的文献求助20
刚刚
dairenfe发布了新的文献求助10
1秒前
Biofly526完成签到,获得积分10
1秒前
2秒前
2秒前
情怀的应助被洋1采纳,获得10
2秒前
如风发布了新的文献求助10
3秒前
3秒前
dcd7777发布了新的文献求助10
3秒前
OTW发布了新的文献求助10
4秒前
云杉木发布了新的文献求助10
4秒前
sjsknd给sjsknd的求助进行了留言
4秒前
脑洞疼的应助被ss毒是采纳,获得10
4秒前
庸俗完成签到,获得积分10
4秒前
FashionBoy的应助被LL采纳,获得10
5秒前
5秒前
蔡胜森完成签到,获得积分10
5秒前
5秒前
李爱国的应助被非我采纳,获得10
6秒前
6秒前
7秒前
向前发布了新的文献求助10
7秒前
诗谙发布了新的文献求助10
7秒前
Jacky完成签到,获得积分10
7秒前
浏阳河发布了新的文献求助10
8秒前
酷波er的应助被wc采纳,获得10
8秒前
懵懂的枫叶完成签到 ,获得积分10
8秒前
刘子完成签到 ,获得积分10
10秒前
科目三的应助被ldy采纳,获得10
10秒前
化学废材完成签到,获得积分10
11秒前
虚拟的如霜完成签到,获得积分20
11秒前
11秒前
希望天下0贩的0的应助被RATHER采纳,获得10
12秒前
优雅的东完成签到,获得积分10
12秒前
JUgu发布了新的文献求助10
12秒前
浏阳河完成签到,获得积分10
12秒前
激情的惜霜完成签到,获得积分10
13秒前
大模型的应助被鲤黎黎采纳,获得10
13秒前
传奇3的应助被sanbuzhiwai采纳,获得30
13秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
Organizational Behavior 510
Arbitrage Theory in Discrete and Continuous Time 500
Fortepian Chopina 400
A Silent Apostrophe:The Fayum Portraits 310
四川大学学位论文.郭瑞昂. 基于高压热扩散的n型磷掺杂金刚石半导体制备研究 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 有机化学 化学工程 内科学 物理 生物化学 复合材料 催化作用 细胞生物学 人工智能 心理学 无机化学 基因 遗传学
热门帖子
关注 科研通微信公众号,转发送积分 7830656
求助须知:如何正确求助?哪些是违规求助? 9355069
关于积分的说明 20582620
捐赠科研通 7423494
什么是DOI,文献DOI怎么找? 3336511
关于科研通互助平台的介绍 2481062
邀请新用户注册赠送积分活动 2357076