清晨好,您是今天最早来到科研通的研友!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您科研之路漫漫前行!

VisualRAG: Knowledge-Guided Retrieval Augmentation for Image-Text Matching

计算机科学 图像检索 人工智能 计算机视觉 图像匹配 图像(数学) 匹配(统计) 模式识别(心理学) 情报检索 数学 统计
作者
Hengchang Wang,Li Liu,Huaxiang Zhang,Lei Zhu,Xiaojun Chang,Hao Du
出处
期刊:IEEE Transactions on Circuits and Systems for Video Technology [Institute of Electrical and Electronics Engineers]
卷期号:36 (1): 1234-1248 被引量:1
标识
DOI:10.1109/tcsvt.2025.3597097
摘要

Image-text matching as a fundamental cross-modal understanding task presents unique challenges in weakly-aligned scenarios. Such data typically feature highly abstract textual captions with sparse entity references, creating a significant semantic gap with visual content. Current mainstream methods, primarily designed for strongly aligned data pairs, employ dynamic modeling or multi-dimensional similarity computation to achieve feature space mapping. However, they struggle with information asymmetry and modal heterogeneity in weakly aligned cases. To address this, we propose a Visual Perception Knowledge Enhancement (VPKE) framework. Unlike existing methods based on strong alignment assumptions, this framework mines latent image semantics through vision-language models and generates auxiliary captions, overcoming the information bottleneck of traditional text modalities. Its core innovation lies in an adaptive knowledge distillation mechanism that combines retrieval-augmented generation (RAG) with key entity extraction. This mechanism effectively filters noise when introducing external knowledge while optimizing cross-modal feature integration. The framework employs multi-level similarity evaluation to dynamically adjust fusion weights among original text, key entities, and auxiliary captions, enabling adaptive integration of diverse semantic features and significantly improving model flexibility. Additionally, multi-scale feature extraction further enhances cross-modal representation capabilities. Experimental results show that the proposed method performs excellently in image-text retrieval tasks on the MSCOCO and Flickr30K datasets, validating its effectiveness.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
沙海沉戈完成签到,获得积分0
28秒前
俊逸吐司完成签到 ,获得积分10
36秒前
小二郎应助白华苍松采纳,获得10
42秒前
彩色樱桃完成签到,获得积分10
43秒前
无悔完成签到 ,获得积分0
45秒前
活泼晓兰完成签到,获得积分20
53秒前
记上没文献了完成签到 ,获得积分10
54秒前
李太白云游四海完成签到,获得积分10
1分钟前
忘忧Aquarius完成签到,获得积分0
1分钟前
cadcae完成签到,获得积分10
1分钟前
害羞的又菡完成签到,获得积分10
1分钟前
1分钟前
年轻花卷完成签到,获得积分10
1分钟前
1分钟前
bkagyin应助迟到翘课翘采纳,获得10
1分钟前
年轻静蕾完成签到,获得积分10
2分钟前
scenery0510完成签到,获得积分0
2分钟前
2分钟前
2分钟前
乐乐应助白华苍松采纳,获得10
2分钟前
2分钟前
MingY完成签到,获得积分10
2分钟前
棉裤完成签到,获得积分10
2分钟前
怕黑明雪完成签到,获得积分10
2分钟前
77完成签到,获得积分10
2分钟前
刘雯完成签到,获得积分10
2分钟前
晨风完成签到,获得积分10
2分钟前
曾经的盼望完成签到,获得积分10
3分钟前
spinon完成签到,获得积分10
3分钟前
紫熊完成签到,获得积分10
3分钟前
平淡的友儿完成签到 ,获得积分10
3分钟前
善良士晋完成签到,获得积分10
3分钟前
白华苍松完成签到,获得积分10
3分钟前
情怀应助白华苍松采纳,获得10
4分钟前
橘子完成签到 ,获得积分10
4分钟前
Axs完成签到,获得积分10
4分钟前
靓丽的小懒虫完成签到,获得积分10
4分钟前
humorlife完成签到,获得积分10
4分钟前
现代的冰海完成签到,获得积分10
4分钟前
zyyicu完成签到,获得积分10
4分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
自動車の空力技術 800
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7778308
求助须知:如何正确求助?哪些是违规求助? 9318775
关于积分的说明 20365875
捐赠科研通 7365397
什么是DOI,文献DOI怎么找? 3319203
关于科研通互助平台的介绍 2467059
邀请新用户注册赠送积分活动 2334591