计算机科学
图像检索
人工智能
对偶(语法数字)
视觉文字
计算机视觉
图像自动标注
图像处理
图像(数学)
模式识别(心理学)
情报检索
数据检索
图像分割
视频检索
文献检索
文档图像处理
特征提取
图像纹理
相位恢复
可视化
基于内容的图像检索
搜索引擎索引
查询扩展
图像压缩
作者
Wenyue Tang,Jianze Wei,Xingyu Gao
标识
DOI:10.1109/tip.2025.3597043
摘要
Composed Image Retrieval (CIR) is a popular multi-modal retrieval task that aims to retrieve a target image based on a query composed of a reference image and modification text. The challenge lies in how to effectively retrieve a target image that preserves the visual content of the reference image while incorporating the changes described by the modification text. Existing CIR methods primarily employ a fusion-based strategy or a textual-inversion strategy during training. Although these methods have achieved promising results, they are limited in fully leveraging multi-modal information. This results in modality redundancy, where the retrieval process is dominated by one modality while ignoring the other. To address this issue, we propose an asymmetric fusion mechanism to generate dual retrieval queries of different granularity, enabling the model to fully use multi-modal information. Specifically, we propose a novel method termed Dual Retrieval Queries Fine-Tuning for Composed Image Retrieval (DRQ-CIR), which consists of two key components: 1) a Bilateral Multi-Modal Fusion (BMMF) module based on pre-trained VLMs, which combines the reference image and modification text to generate an enriched retrieval query; and 2) a Dual Retrieval Queries Fine-Tuning (DRQ-FT) module, which employs latent prompts to generate an enhanced retrieval query. Dual retrieval queries are used for contrastive learning with the target image to fine-tune the model and improve retrieval performance. Additionally, we introduce a bi-directional training paradigm to ensure retrieval consistency and further exploit the triplets. Extensive experiments validate the effectiveness of our proposed method on four established benchmark datasets. (Code will be available at: https://github.com/Crystal-twy998/DRQ-CIR).
科研通智能强力驱动
Strongly Powered by AbleSci AI