Can AI speak endo? A multi-platform evaluation of large language models against ESHRE endometriosis guidelines

指南 一致性(知识库) 子宫内膜异位症 医学 可靠性(半导体) 计算机科学 医学物理学 语言模型 梅德林 自然语言处理 妇科 情报检索 临床实习 循证医学 人工智能 短信 Web应用程序 医学文献 医学诊断 可用性 家庭医学 机器学习
作者
Gaetano Riemma,Florindo Mario Caniglia,Camilla Casolari,Antonio Maiorana,Mauro Cozzolino,Vittorio Agrifoglio,Pasquale De Franciscis,Luigi Cobellis,Raffaela Carotenuto,Andrea Etrusco
出处
期刊:Human Reproduction [Oxford University Press]
标识
DOI:10.1093/humrep/deag116
摘要

STUDY QUESTION: How do the three most widely accessible large language models perform in terms of accuracy, consistency, and reliability when answering clinically relevant questions derived from the 2022 ESHRE guideline on endometriosis? SUMMARY ANSWER: Model A achieved the highest accuracy scores, model C demonstrated significantly superior consistency across repeated queries, and all three models showed comparable but suboptimal reliability under free-tier access, while under official API access median accuracy converged across providers and reliability increased. WHAT IS KNOWN ALREADY: Large language models are increasingly consulted by both clinicians and patients as readily accessible sources of medical information. In reproductive medicine, preliminary evidence suggests that individual platforms may retrieve endometriosis-related content with mixed fidelity. However, no study has simultaneously benchmarked multiple models against a single, internationally recognized endometriosis guideline. The extent to which these tools can be trusted to faithfully reproduce evidence-based recommendations on endometriosis diagnosis and management remains largely unexplored. STUDY DESIGN, SIZE, DURATION: Cross-sectional, multi-platform comparative study. Fifty clinically relevant questions covering the diagnostic and therapeutic domains of the 2022 ESHRE endometriosis guideline were simultaneously submitted to all three models during December 2025. Each question was entered in duplicate using independent sessions to assess response consistency and reliability. Subgroup analyses were carried out by submitting the questions to the API version and to the free-tier version of the platforms available in May 2026. PARTICIPANTS/MATERIALS, SETTING, METHODS: The three models were accessed through their respective free web interfaces using new accounts, without prompt engineering, prior training, or retrieval-augmented generation. A zero-shot prompting approach was adopted. Accuracy was evaluated by two independent experts using the Global Quality Score (GQS); consistency was defined as identical responses across the three iterations; reliability was defined as the alignment of each response with the ESHRE guideline. Discrepancies were settled by a third reviewer. MAIN RESULTS AND THE ROLE OF CHANCE: Significant differences in accuracy were observed across models (Kruskal-Wallis H = 37.10, P < 0.001). Model A achieved the highest median GQS (5, interquartile range [IQR] 4-5), followed by model C (4, IQR 3-5) and model B (3, IQR 3-4). Post-hoc analysis confirmed that model A significantly outperformed model B (P < 0.001) but not model C (P = 0.123), while model C also scored significantly higher than model B (P < 0.001). For consistency, model C demonstrated a significantly higher rate of reproducible responses (92.0%) compared with model A (72.0%, P = 0.028) and model B (68.0%, P = 0.008). No significant between-model differences were found for reliability (χ2 = 1.029, P = 0.598), with rates of 76.0% for model C, 68.0% for model A, and 68.0% for model B. In API and free-tier May 2026 subgroup analyses, median GQS converged to 4 across all three providers, and reliability rose to at least 76% for each provider. LIMITATIONS, REASONS FOR CAUTION: Model outputs may change with subsequent updates. GQS retains a degree of subjectivity despite expert adjudication. The study tested factual recall rather than complex clinical reasoning, limiting generalizability to real-world decision-making scenarios. WIDER IMPLICATIONS OF THE FINDINGS: These findings provide the first multi-platform benchmark of large language models against the latest endometriosis guideline. While models A and C retrieved guideline-concordant information with acceptable fidelity, none of the models achieved a level of reliability required for unsupervised clinical use. The dissociation between accuracy and consistency at the free tier, and its attenuation under API and more recent accesses, indicates that both model capability and commercial tier shape user-facing outputs. The dissociation between accuracy and consistency underscores that a model producing high-quality answers does not necessarily do so in a reproducible manner. Expert oversight remains crucial when interpreting their output regarding endometriosis care. Future research should extend this framework to additional endometriosis guidelines and to longitudinal monitoring of models' performance with their upgrades. FUNDING: No external funding was received for this study. DISCLOSURES: The authors declare no competing interests. TRIAL REGISTRATION NUMBER: N/A.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
西北望发布了新的文献求助10
刚刚
康德完成签到,获得积分10
刚刚
薛吒发布了新的文献求助10
2秒前
Joe完成签到,获得积分10
2秒前
2秒前
豆包好友发布了新的文献求助10
2秒前
cc完成签到,获得积分10
3秒前
Hzw398发布了新的文献求助10
3秒前
3秒前
3秒前
3秒前
3秒前
GraceWu完成签到,获得积分10
4秒前
4秒前
5秒前
6秒前
Blue发布了新的文献求助10
6秒前
香蕉觅云应助虎攀伟采纳,获得10
7秒前
可爱的函函应助浮生绘采纳,获得10
7秒前
8秒前
CodeCraft应助清爽鼠标采纳,获得10
8秒前
Kumiko发布了新的文献求助10
9秒前
神勇的博涛完成签到,获得积分10
9秒前
9秒前
2421154880完成签到,获得积分10
10秒前
心肌boy完成签到,获得积分20
10秒前
10秒前
dlCao发布了新的文献求助10
11秒前
Q97发布了新的文献求助10
13秒前
rainer完成签到,获得积分10
13秒前
jstagey完成签到,获得积分10
14秒前
烨霖完成签到,获得积分10
14秒前
14秒前
爆米花应助CC采纳,获得10
16秒前
小葡萄完成签到,获得积分10
16秒前
Jiling发布了新的文献求助10
16秒前
自然寒荷应助何时采纳,获得10
17秒前
爆米花应助张张采纳,获得10
17秒前
科目三应助悦悦采纳,获得10
17秒前
冷酷保温杯完成签到,获得积分10
17秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Effects of Two Weeks of Red Light Therapy on Choroidal Thickness and Axial Length in Young Adults 700
内視鏡的に摘除しえた十二指腸乳頭部腫瘍の2例 660
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The Neuroscience of Language 400
Common Foundations of American and East Asian Modernisation: From Alexander Hamilton to Junichero Koizumi 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7675198
求助须知:如何正确求助?哪些是违规求助? 9241491
关于积分的说明 19911816
捐赠科研通 7245075
什么是DOI,文献DOI怎么找? 3286117
关于科研通互助平台的介绍 2444163
邀请新用户注册赠送积分活动 2288550