Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval

计算机科学 简单(哲学) 原始数据 图像融合 人工智能 融合 传感器融合 图像检索 图像(数学) 计算机视觉 模式识别(心理学) 情报检索 哲学 语言学 认识论 程序设计语言
作者
Haokun Wen,Xuemeng Song,Xiaolin Chen,Yinwei Wei,Liqiang Nie,Tat‐Seng Chua
标识
DOI:10.1145/3626772.3657727
摘要

Composed image retrieval (CIR) aims to retrieve the target image based on a multimodal query, i.e., a reference image paired with corresponding modification text. Recent CIR studies leverage vision-language pre-trained (VLP) methods as the feature extraction backbone, and perform nonlinear feature-level multimodal query fusion to retrieve the target image. Despite the promising performance, we argue that their nonlinear feature-level multimodal fusion may lead to the fused feature deviating from the original embedding space, potentially hurting the retrieval performance. To address this issue, in this work, we propose shifting the multimodal fusion from the feature level to the raw-data level to fully exploit the VLP model's multimodal encoding and cross-modal alignment abilities. In particular, we introduce a Dual Query Unification-based Composed Image Retrieval framework (DQU-CIR), whose backbone simply involves a VLP model's image encoder and a text encoder. Specifically, DQU-CIR first employs two training-free query unification components: text-oriented query unification and vision-oriented query unification, to derive a unified textual and visual query based on the raw data of the multimodal query, respectively. The unified textual query is derived by concatenating the modification text with the extracted reference image's textual description, while the unified visual query is created by writing the key modification words onto the reference image. Ultimately, to address diverse search intentions, DQU-CIR linearly combines the features of the two unified queries encoded by the VLP model to retrieve the target image. Extensive experiments on four real-world datasets validate the effectiveness of our proposed method.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
duckspy发布了新的文献求助10
刚刚
caicai发布了新的文献求助10
1秒前
1秒前
1秒前
2秒前
温暖月饼发布了新的文献求助10
2秒前
小恐龙发布了新的文献求助10
2秒前
忐忑的芸发布了新的文献求助10
3秒前
Orange应助睡不醒的酸奶采纳,获得10
4秒前
4秒前
小小蝶完成签到,获得积分20
5秒前
5秒前
5秒前
6秒前
小小陈完成签到,获得积分10
6秒前
caicai完成签到,获得积分10
6秒前
益善完成签到,获得积分10
6秒前
16680018995完成签到,获得积分20
7秒前
8秒前
你好发布了新的文献求助10
8秒前
彩虹糖给彩虹糖的求助进行了留言
9秒前
任娜发布了新的文献求助10
9秒前
秋风发布了新的文献求助10
9秒前
喷火娃应助guagua采纳,获得10
9秒前
9秒前
9秒前
郭郭完成签到 ,获得积分10
9秒前
文静犀牛完成签到,获得积分10
9秒前
xia xianxin完成签到,获得积分10
10秒前
10秒前
菱歌万金发布了新的文献求助10
11秒前
11秒前
宣登仕发布了新的文献求助10
11秒前
DBT完成签到,获得积分10
12秒前
12秒前
哄小孩的广君完成签到,获得积分20
13秒前
13秒前
13秒前
思源应助shotaro采纳,获得10
14秒前
历飞雨完成签到 ,获得积分10
14秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
日本現代怪異事典 副読本 700
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 650
Machine Learning for Asset Management and Pricing 600
Numerical analysis of the coupled atmosphere-ocean models (CAO II). II 600
Models for the coupled atmosphere and ocean 600
Évora na Idade Média 555
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7382344
求助须知:如何正确求助?哪些是违规求助? 8989571
关于积分的说明 19122338
捐赠科研通 7021195
什么是DOI,文献DOI怎么找? 3227172
关于科研通互助平台的介绍 2390203
邀请新用户注册赠送积分活动 2208038