Evaluating reasoning in multimodal large language models for ophthalmology: a bilingual benchmark study using clinical vignettes and imaging

印为红字的 医学 渐晕 人工智能 自然语言处理 可解释性 机器学习 稳健性(进化) 医学教育 梅德林 水准点(测量) 子专业 计算机科学 医学影像学 医学物理学 医学诊断 复杂度 理解力 模式治疗法 客观结构化临床检查 定性性质 质量(理念)
作者
Houfa Yin,Kaikai Zhao,Danli Shi,Andrzej Grzybowski,Kai Jin
出处
期刊:British Journal of Ophthalmology [BMJ]
卷期号:: bjo-2025
标识
DOI:10.1136/bjo-2025-328992
摘要

BACKGROUND: Large language models (LLMs) excel in text-based medical exams, but their ability to integrate multimodal data, critical for ophthalmology, is underexplored. This study evaluates vision-language LLMs' accuracy and reasoning in complex ophthalmic questions. METHODS: We assessed three multimodal LLMs (CLM-V, ChatGPT-5, MiniCPM-V 4.5) on 316 bilingual ophthalmology questions (175 English Basic and Clinical Science Course single-choice, 141 Chinese senior professional title multiple-choice questions) across cornea, uvea, glaucoma, retina and orbit. Each question paired a clinical vignette with an image. Models were tested with reasoning-enabled and reasoning-disabled prompts. Accuracy was measured against reference standards, and reasoning quality was evaluated using automated rubric scoring (accuracy, data synthesis, logic, option analysis, safety) and expert review. Four cases were analysed qualitatively. RESULTS: Reasoning-enabled prompting increased mean artificial intelligence-assisted total scores in the English dataset from 14.97 to 16.07 for CLM-V, 20.77 to 23.97 for ChatGPT-5 and 10.83 to 12.60 for MiniCPM-V 4.5; in the Chinese dataset, the corresponding scores were 9.03 to 10.27, 19.95 to 22.00 and 11.05 to 13.30, respectively. Human evaluation showed substantial inter-rater agreement (κ=0.87) and again ranked ChatGPT-5 highest. Qualitative case analyses illustrated that reasoning-enabled outputs were often more clinically interpretable, although the magnitude of benefit was model-dependent and dataset-dependent. CONCLUSION: Multimodal LLMs demonstrate potential in ophthalmic question-answering, with reasoning-enabled prompting being associated with improved interpretability and, in most settings, numerically higher performance. However, limitations in subspecialty robustness and image interpretation necessitate rigorous reasoning evaluation for safe educational and clinical applications.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
Nole应助uraylong采纳,获得30
刚刚
cotton_04完成签到,获得积分10
刚刚
活着完成签到 ,获得积分10
刚刚
冯不疯完成签到,获得积分10
刚刚
空青完成签到,获得积分10
1秒前
1秒前
FYJ发布了新的文献求助10
1秒前
1秒前
阿高完成签到,获得积分20
1秒前
1秒前
earthclean发布了新的文献求助10
2秒前
开放的从菡完成签到 ,获得积分10
3秒前
3秒前
Shawn完成签到,获得积分10
3秒前
jiedaocheng完成签到,获得积分10
3秒前
学术文献互助应助蜀安采纳,获得200
4秒前
fairy发布了新的文献求助10
4秒前
顾矜应助小张采纳,获得10
4秒前
落幕熊猫完成签到,获得积分0
4秒前
江上浩月完成签到,获得积分10
5秒前
陶醉的大炮完成签到,获得积分10
5秒前
诚心的志泽完成签到,获得积分10
5秒前
5秒前
文艺的熠彤完成签到,获得积分10
6秒前
呱呱完成签到,获得积分10
6秒前
空青发布了新的文献求助10
6秒前
认真芷容完成签到,获得积分10
6秒前
Honey发布了新的文献求助10
6秒前
7秒前
keten完成签到,获得积分10
7秒前
7秒前
7秒前
7秒前
7秒前
7秒前
riccixuu完成签到 ,获得积分10
7秒前
8秒前
twss发布了新的文献求助10
8秒前
李皮皮完成签到 ,获得积分10
8秒前
orixero应助PositiveJugend采纳,获得10
8秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
The anomeric effect 1000
Principles of town planning: translating concepts to applications 1000
Navigating Normative Orders: Interdisciplinary Perspectives 750
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7733078
求助须知:如何正确求助?哪些是违规求助? 9283945
关于积分的说明 20161671
捐赠科研通 7310903
什么是DOI,文献DOI怎么找? 3304251
关于科研通互助平台的介绍 2457078
邀请新用户注册赠送积分活动 2313480