Chatbots in urology: accuracy, calibration, and comprehensibility; is DeepSeek taking over the throne?

可读性 校准 克朗巴赫阿尔法 利克特量表 医学物理学 科恩卡帕 卡帕 比例(比率) 可用性 医学 计算机科学 心理学 统计 临床心理学 心理测量学 数学 机器学习 发展心理学 人机交互 物理 几何学 量子力学 程序设计语言
作者
Ömer Faruk Asker,Muhammed Selim Recai,Yunus Emre Genç,Kader Ada Dogan,Tarık Emre Şener,Bahadır Şahin
出处
期刊:BJUI [Wiley]
卷期号:136 (5): 937-945 被引量:4
标识
DOI:10.1111/bju.16873
摘要

OBJECTIVE: To evaluate widely used chatbots' accuracy, calibration error, readability, and understandability with objective measurements by 35 questions derived from urology in-service examinations, as the integration of large language models (LLMs) into healthcare has gained increasing attention, raising questions about their applications and limitations. MATERIALS AND METHODS: A total of 35 European Board of Urology questions were asked to five LLMs with a standardised prompt that was systematically designed and used across all models: ChatGPT-4o, DeepSeek-R1, Gemini, Grok-2, and Claude 3.5. Accuracy was calculated by Cohen's kappa for all models. Readability was assessed by Flesch Reading Ease, Gunning Fog, Coleman-Liau, Simple Measure of Gobbledygook, and Automated Readability Index, while understandability was determined by scores of residents' ratings by a Likert scale. RESULTS: The models and answer key were in substantial agreement with a Fleiss' kappa of 0.701, and Cronbach's alpha of 0.914. For accuracy, Cohen's kappa was 0.767 for ChatGPT-4o, 0.764 for DeepSeek-R, and 0.765 for Grok-2 (80% accuracy for each), followed by 0.729 for Claude 3.5 (77% accuracy) and 0.611 for Gemini (68.4% accuracy). The lowest calibration error was found in ChatGPT-4o (19.2%) and DeepSeek-R1 scored the highest for readability. In understandability analysis, Claude 3.5 had the highest rating compared to others. CONCLUSION: Chatbots demonstrated various powers across different tasks. DeepSeek-R1, despite being just released, showed promising results in medical applications. These findings highlight the need for further optimisation to better understand the applications of chatbots in urology.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
迷人兔子给迷人兔子的求助进行了留言
刚刚
smiles完成签到,获得积分10
1秒前
桐桐应助传统的天蓝采纳,获得10
1秒前
1秒前
酷波er应助peng采纳,获得10
2秒前
贪玩夏柳完成签到,获得积分10
2秒前
今后应助任性的思远采纳,获得10
2秒前
allrubbish发布了新的文献求助10
2秒前
星辰大海应助原野采纳,获得10
2秒前
Rabbit发布了新的文献求助10
2秒前
科研通AI6.4应助y9gyn_37采纳,获得10
3秒前
3秒前
3秒前
4秒前
4秒前
天天快乐应助蓝朱采纳,获得10
4秒前
今后应助lll采纳,获得10
5秒前
6秒前
幽默跳跳糖完成签到 ,获得积分10
6秒前
7秒前
无花果应助19863737023采纳,获得10
7秒前
Alone完成签到,获得积分10
7秒前
OO发布了新的文献求助10
8秒前
科研通AI6.4应助嘎嘎嘎嘎采纳,获得30
8秒前
慕青应助坦率采纳,获得10
9秒前
光锥之外发布了新的文献求助10
10秒前
小鹿完成签到,获得积分10
11秒前
Alone发布了新的文献求助10
11秒前
核桃发布了新的文献求助10
12秒前
li完成签到,获得积分20
13秒前
13秒前
AireenBeryl531应助科研小狗采纳,获得10
14秒前
14秒前
Duke完成签到,获得积分10
14秒前
贾贡献应助桑梓采纳,获得10
14秒前
轻轻的松松完成签到,获得积分10
15秒前
Nosevaya完成签到 ,获得积分10
15秒前
山楂完成签到,获得积分20
15秒前
NexusExplorer应助leez采纳,获得10
16秒前
SciGPT应助任性的思远采纳,获得10
16秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
内視鏡的に摘除しえた十二指腸乳頭部腫瘍の2例 660
On nonlinear stability of contact discontinuities. In: Hyperbolic problems: theory, numerics, applications (Stony Brook, NY, 1994) 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
微电子器件实验教程 400
The Neuroscience of Language 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7678785
求助须知:如何正确求助?哪些是违规求助? 9243992
关于积分的说明 19926997
捐赠科研通 7249679
什么是DOI,文献DOI怎么找? 3287252
关于科研通互助平台的介绍 2444982
邀请新用户注册赠送积分活动 2290462