亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

Beyond Static Knowledge: Dynamic Context-Aware Cross-Modal Contrastive Learning for Medical Visual Question Answering

计算机科学 答疑 人工智能 工作流程 语义学(计算机科学) 自然语言处理 特征(语言学) 自然语言 代表(政治) 编码(内存) 可视化 模态(人机交互) 知识表示与推理 特征学习 滤波器(信号处理) 适应(眼睛) 感知 机器学习 抽象 医学影像学 钥匙(锁) 编码(集合论) 语义记忆 对比度(视觉) 视觉感受 源代码 基于知识的系统 医学诊断 视觉推理 情报检索 噪音(视频) 特征提取 人机交互 透视图(图形)
作者
Rui Yang,Lijun Liu,Xupeng Feng,Wei Peng,Xiaobing Yang
出处
期刊:IEEE Transactions on Medical Imaging [Institute of Electrical and Electronics Engineers]
卷期号:45 (3): 1075-1087
标识
DOI:10.1109/tmi.2025.3617289
摘要

Medical Visual Question Answering (Med-VQA) aims to analyze medical images and accurately respond to natural language queries, thereby optimizing clinical workflows and improving diagnostic and therapeutic outcomes. Although medical images contain rich visual information, the corresponding textual queries frequently lack sufficient descriptive content. This imbalance of information and modality differences leads to significant semantic bias. Furthermore, existing approaches integrate external medical knowledge to enhance model performance, they primarily rely on static knowledge that lacks dynamic adaptation to specific input samples, leading to redundant information and noise interference. To address these challenges, we propose a Contextual Knowledge-Aware Dynamic Perception for the Cross-Modal Reasoning and Alignment (CKRA) Model. To mitigate knowledge redundancy, CKRA employs a dynamic perception mechanism that leverages semantic cues from the query to selectively filter relevant medical knowledge specific to the current sample's context. To alleviate cross-modal semantic bias, CKRA bridges the distance between visual and linguistic features through knowledge-image contrastive learning, optimizing knowledge feature representation and directing the model's attention to key image regions. Further, we design a dual-stream guided attention network that facilitates cross-modal interaction and alignment across multiple dimensions. Experimental results show that the proposed CKRA model outperforms the state-of-the-art method on SLAKE and VQA-RAD datasets. In addition, ablation studies validate the effectiveness of each module, while Grad-CAM maps further demonstrate the feasibility of CKRA for medical visual questioning tasks. The source code and weights of the model are available at https://github.com/cloneiq/CKRA-MedVQA.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
2秒前
沭阳检验医师完成签到,获得积分0
4秒前
yat发布了新的文献求助30
5秒前
17秒前
20秒前
洁净友蕊完成签到,获得积分10
23秒前
活泼画板发布了新的文献求助10
26秒前
AliEmbark发布了新的文献求助10
26秒前
周一更发布了新的文献求助10
52秒前
海绵宝宝完成签到 ,获得积分10
56秒前
奋斗不二完成签到,获得积分10
1分钟前
研友_LMo56Z完成签到,获得积分10
1分钟前
迷人白桃完成签到,获得积分10
2分钟前
2分钟前
2分钟前
chethiran发布了新的文献求助10
2分钟前
丘比特应助进击的咩咩采纳,获得50
2分钟前
Criminology34举报稳赚赚求助涉嫌违规
2分钟前
魁梧的天佑完成签到,获得积分10
2分钟前
2分钟前
3分钟前
蛋白积聚完成签到,获得积分10
3分钟前
3分钟前
NexusExplorer应助Snow886采纳,获得30
3分钟前
耍酷笑白发布了新的文献求助10
3分钟前
3分钟前
3分钟前
XXshen发布了新的文献求助10
3分钟前
Snow886发布了新的文献求助30
3分钟前
3分钟前
科研通AI6.4应助yat采纳,获得10
3分钟前
学生信的大叔完成签到,获得积分10
3分钟前
搜集达人应助耍酷笑白采纳,获得10
3分钟前
周一更发布了新的文献求助10
3分钟前
英俊的谷蓝完成签到,获得积分10
3分钟前
3分钟前
洗月完成签到 ,获得积分10
3分钟前
3分钟前
3分钟前
Snow886发布了新的文献求助30
3分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Römisch-Germanische Forschungen 1000
APA handbook of comparative psychology: Basic concepts, methods, neural substrate, and behavior 1000
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
Electric machines: theory, operating applications, and controls 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7605169
求助须知:如何正确求助?哪些是违规求助? 9181029
关于积分的说明 19662329
捐赠科研通 7179854
什么是DOI,文献DOI怎么找? 3269491
关于科研通互助平台的介绍 2433424
邀请新用户注册赠送积分活动 2263580