Large language models for risk-of-bias assessment in randomised clinical trials—a comparative validation study

计算机科学 自然语言处理 梅德林 医学 人工智能 医学物理学 心理学 模型验证 临床试验 数据科学 研究设计 机器学习
作者
Lauri Nyrhi,Ville Ponkilainen,Juho Laaksonen,Lauri Kuikka,Lauri Paljakka,Teemu Karjalainen,Ville M. Mattila,Ilari Kuitunen
出处
期刊:EBioMedicine [Elsevier BV]
卷期号:126: 106238-106238 被引量:2
标识
DOI:10.1016/j.ebiom.2026.106238
摘要

Background Large language models (LLMs) are emerging tools for evidence synthesis. Risk of bias (RoB) assessment of trials remains an essential but time-consuming step inconsistent even amongst experts. Early LLM studies showed mixed reliability. Advances in reasoning-enabled models warrant evaluation of their accuracy and consistency for RoB screening across randomised trials to reduce reviewer workload. Methods We conducted a preregistered comparative validation study (March 11–May 19, 2025) of four LLMs—ChatGPT o3, DeepSeek v3, Google Gemini Flash 2.0, and Grok 3—prompted with full-text randomised clinical trial articles and protocols. Two corpora were analysed: 100 RCTs from recent Cochrane reviews (RoB 1) and 100 RCTs from meta-analyses in high-impact journals (RoB 2). The reference standard was published human RoB judgements. The primary outcome was interobserver reliability (Cohen κ, 95% CI); secondary outcomes were intraobserver agreement and diagnostic accuracy (sensitivity, specificity, predictive values, F 1 -score). Findings For RoB 1, interobserver agreement ranged from κ 0.0.27 (95% CI 0.07–0.46) with Gemini Flash 2.0 to κ 0.39 (0.20–0.59) with DeepSeek v3. For RoB 2, agreement was lower, from κ 0.06 (−0.07 to 0.18) with ChatGPT o3 to κ 0.13 (−0.04 to 0.31) with Gemini. Diagnostic performance was limited with sensitivity ranging 0.05–0.55, specificity 0.78–0.99, PPV 0.31–0.50, and NPV 0.48–0.61 across models, with models consistently over-flagging concerns. Interpretation None of the evaluated LLMs were sufficiently reliable for fully autonomous RoB assessment. DeepSeek v3 and ChatGPT o3 approximated human performance best on RoB 1, but RoB 2 rule-in and rule-out performance remained modest. Current use should be supervised, with possible application of LLMs for triage or as a second assessor. Major improvements in protocol retrieval, task-specific tuning, and calibrated thresholds, prospectively validated, are needed for safe stand-alone deployment. Funding This study received no financial support.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
火星上的小笼包完成签到,获得积分10
刚刚
ATREE发布了新的文献求助10
刚刚
嗨皮尼斯发布了新的文献求助10
1秒前
ly发布了新的文献求助10
1秒前
1秒前
orixero应助xg_kim采纳,获得10
1秒前
2秒前
12345678发布了新的文献求助10
2秒前
3秒前
3秒前
3秒前
woshi123应助乘风采纳,获得10
4秒前
LittleTT发布了新的文献求助10
4秒前
4秒前
CQUzc完成签到,获得积分10
4秒前
AEB完成签到,获得积分10
5秒前
故渊丶发布了新的文献求助10
5秒前
6秒前
科研通AI6.2应助1820采纳,获得10
6秒前
草雨田完成签到,获得积分10
6秒前
聪明的身影完成签到,获得积分10
7秒前
7秒前
拉长的诗蕊完成签到,获得积分10
7秒前
汉堡包应助junjun采纳,获得10
7秒前
JamesPei应助8023采纳,获得10
8秒前
qiqi完成签到,获得积分10
8秒前
chao发布了新的文献求助10
8秒前
kang发布了新的文献求助10
8秒前
小马甲应助楚天采纳,获得10
9秒前
Function完成签到,获得积分10
9秒前
邪灬坤发布了新的文献求助10
10秒前
10秒前
张zy完成签到,获得积分10
10秒前
yangjun发布了新的文献求助10
11秒前
传奇3应助嘟嘟巴拉巴拉采纳,获得10
11秒前
Nan发布了新的文献求助10
12秒前
Seameng完成签到 ,获得积分10
12秒前
12秒前
斯文败类应助wwwq采纳,获得10
12秒前
小二郎应助Augenstern采纳,获得10
13秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
An Introduction to Foreign Language Learning and Teaching 750
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
Synthesis of P-Chiral Phosphine Ligands and Their Applications in Asymmetric Catalysis 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7622529
求助须知:如何正确求助?哪些是违规求助? 9197835
关于积分的说明 19716458
捐赠科研通 7194042
什么是DOI,文献DOI怎么找? 3272988
关于科研通互助平台的介绍 2435430
邀请新用户注册赠送积分活动 2268373