计算机科学
答疑
一般化
限制
人工智能
自然语言处理
钥匙(锁)
语言模型
机器学习
编码(集合论)
可视化
语义学(计算机科学)
情报检索
编码(内存)
视觉语言
统一医学语言系统
自然语言
医学影像学
对比度(视觉)
人机交互
作者
Zefan Zhang,Yanhui Li,Ruihong Zhao,Tian Bai
标识
DOI:10.1109/jbhi.2026.3663420
摘要
Difference-aware Medical Visual Question Answering (MVQA) aims to answer questions regarding disease-related content and the visual differences between the paired medical images, which is crucial for assessing disease progression and guiding further treatment planning. Although current medical Multimodal Large Language Models (MLLMs) have shown promising results in MVQA, they still exhibit poor generalization performance in difference-aware MVQA due to two key challenges. Firstly, existing difference-aware MVQA datasets are biased toward temporal variations of individual diseases, limiting their ability to model multi-disease coexistence and overlapping symptoms in real-world clinical scenarios. Secondly, disease-level semantic alignment becomes more challenging with multi-image inputs, as they introduce more redundant and interfering visual features. To address the first challenge, we introduce DAMON-QA, a large-scale difference-aware MVQA dataset designed to support visual difference analysis across multiple diseases. Leveraging this dataset, we train MLLMs and propose a Difference-Aware Medical visual questiON answering (DAMON) model. To tackle the second challenge, we further propose a Disease-driven Prompt Module (DPM) to identify the relevant diseases and guide the disease difference analysis process. Experiments on MIMIC-Diff-VQA show that our DAMON model achieves state-of-the-art (SOTA) performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI