桥接(联网)
答疑
计算机科学
语义鸿沟
自然语言处理
人工智能
情报检索
图像(数学)
图像检索
计算机网络
作者
Zilin Lu,Qingjie Zeng,Mengkang Lu,Geng Chen,Yong Xia
标识
DOI:10.1109/tmi.2025.3580561
摘要
Medical Visual Question Answering (Med-VQA) aims to answer questions regarding the content of medical images, crucial for enhancing diagnostics and education in healthcare. However, progress in this field is hindered by data scarcity due to the resource-intensive nature of medical data annotation. While existing Med-VQA approaches often rely on pre-training to mitigate this issue, bridging the semantic gap between pre-trained models and specific tasks remains a significant challenge. This paper presents the Dynamic Semantic-Adaptive Prompting (DSAP) framework, leveraging prompt learning to enhance model performance in Med-VQA. To this end, we introduce two prompting strategies: Semantic Alignment Prompting (SAP) and Dynamic Question-Aware Prompting (DQAP). SAP prompts multi-modal inputs during fine-tuning, reducing the semantic gap by aligning model outputs with domain-specific contexts. Simultaneously, DQAP enhances answer selection by leveraging grammatical relationships between questions and answers, thereby improving accuracy and relevance. The DSAP framework was pre-trained on three datasets-ROCO, MedICaT, and MIMIC-CXR-and comprehensively evaluated against 15 existing Med-VQA models on three public datasets: VQA-RAD, SLAKE, and PathVQA. Our results demonstrate a substantial performance improvement, with DSAP achieving a 1.9% enhancement in average results across benchmarks. These findings underscore DSAP's effectiveness in addressing critical challenges in Med-VQA and suggest promising avenues for future developments in medical AI.
科研通智能强力驱动
Strongly Powered by AbleSci AI