Performance of ChatGPT on the Chinese Postgraduate Examination for Clinical Medicine: Survey Study

医学教育 医学 家庭医学 中医药 传统医学 心理学 替代医学 病理
作者
Peng Yu,Changchang Fang,Xiaolin Liu,Wanying Fu,Jitao Ling,Zhiwei Yan,Yuan Jiang,Zhengyu Cao,Maoxiong Wu,Zhiteng Chen,Wengen Zhu,Yuling Zhang,Ayiguli Abudukeremu,Yue Wang,Xiao Liu,J. Wang
出处
期刊:JMIR medical education [JMIR Publications]
卷期号:10: e48514-e48514 被引量:14
标识
DOI:10.2196/48514
摘要

Background ChatGPT, an artificial intelligence (AI) based on large-scale language models, has sparked interest in the field of health care. Nonetheless, the capabilities of AI in text comprehension and generation are constrained by the quality and volume of available training data for a specific language, and the performance of AI across different languages requires further investigation. While AI harbors substantial potential in medicine, it is imperative to tackle challenges such as the formulation of clinical care standards; facilitating cultural transitions in medical education and practice; and managing ethical issues including data privacy, consent, and bias. Objective The study aimed to evaluate ChatGPT’s performance in processing Chinese Postgraduate Examination for Clinical Medicine questions, assess its clinical reasoning ability, investigate potential limitations with the Chinese language, and explore its potential as a valuable tool for medical professionals in the Chinese context. Methods A data set of Chinese Postgraduate Examination for Clinical Medicine questions was used to assess the effectiveness of ChatGPT’s (version 3.5) medical knowledge in the Chinese language, which has a data set of 165 medical questions that were divided into three categories: (1) common questions (n=90) assessing basic medical knowledge, (2) case analysis questions (n=45) focusing on clinical decision-making through patient case evaluations, and (3) multichoice questions (n=30) requiring the selection of multiple correct answers. First of all, we assessed whether ChatGPT could meet the stringent cutoff score defined by the government agency, which requires a performance within the top 20% of candidates. Additionally, in our evaluation of ChatGPT’s performance on both original and encoded medical questions, 3 primary indicators were used: accuracy, concordance (which validates the answer), and the frequency of insights. Results Our evaluation revealed that ChatGPT scored 153.5 out of 300 for original questions in Chinese, which signifies the minimum score set to ensure that at least 20% more candidates pass than the enrollment quota. However, ChatGPT had low accuracy in answering open-ended medical questions, with only 31.5% total accuracy. The accuracy for common questions, multichoice questions, and case analysis questions was 42%, 37%, and 17%, respectively. ChatGPT achieved a 90% concordance across all questions. Among correct responses, the concordance was 100%, significantly exceeding that of incorrect responses (n=57, 50%; P<.001). ChatGPT provided innovative insights for 80% (n=132) of all questions, with an average of 2.95 insights per accurate response. Conclusions Although ChatGPT surpassed the passing threshold for the Chinese Postgraduate Examination for Clinical Medicine, its performance in answering open-ended medical questions was suboptimal. Nonetheless, ChatGPT exhibited high internal concordance and the ability to generate multiple insights in the Chinese language. Future research should investigate the language-based discrepancies in ChatGPT’s performance within the health care context.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
leapper完成签到 ,获得积分10
1秒前
Charming完成签到,获得积分10
1秒前
锦沫完成签到 ,获得积分10
2秒前
3秒前
跳跃发布了新的文献求助10
3秒前
嘤嘤嘤完成签到 ,获得积分10
3秒前
犹豫的大碗应助99采纳,获得10
3秒前
4秒前
叶落滴滴哒哒完成签到,获得积分10
4秒前
陆陆完成签到 ,获得积分10
4秒前
dearwu发布了新的文献求助10
6秒前
纯真的元风完成签到,获得积分10
6秒前
Sweet完成签到 ,获得积分10
6秒前
menghongmei完成签到 ,获得积分10
7秒前
俞安珊完成签到,获得积分10
7秒前
小龙完成签到,获得积分10
7秒前
寒冷班完成签到,获得积分10
8秒前
8秒前
13799109861发布了新的文献求助10
8秒前
夕荀完成签到,获得积分10
8秒前
智慧完成签到,获得积分10
10秒前
发发旦旦完成签到,获得积分10
11秒前
lyb1853完成签到 ,获得积分10
12秒前
AlphaAvery完成签到 ,获得积分10
12秒前
研友_Z7Xdl8完成签到,获得积分10
13秒前
AHMZI完成签到,获得积分10
14秒前
15秒前
典雅雅容完成签到,获得积分10
15秒前
lyf完成签到,获得积分10
17秒前
顾矜应助kexue采纳,获得10
17秒前
研友_Z7Xdl8发布了新的文献求助30
18秒前
zhangyiyang完成签到 ,获得积分10
18秒前
yjia完成签到 ,获得积分10
19秒前
19秒前
hy完成签到 ,获得积分10
20秒前
碧蓝的幻悲完成签到 ,获得积分10
21秒前
cpufigo发布了新的文献求助10
21秒前
23秒前
23秒前
虚心的飞松完成签到 ,获得积分10
23秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The role of consumer psychology in the marketing strategies of pop mart in Thailand 500
核安全综合知识2024版 500
Photothermal Science and Techniques 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7720888
求助须知:如何正确求助?哪些是违规求助? 9274185
关于积分的说明 20101351
捐赠科研通 7297136
什么是DOI,文献DOI怎么找? 3300271
关于科研通互助平台的介绍 2454172
邀请新用户注册赠送积分活动 2307724