计算机科学
情报检索
图像检索
图像(数学)
计算机视觉
万维网
数据科学
作者
Longbiao Du,S. Deng,Ying Li,Jun Li,Qi Tian
摘要
Composed Image Retrieval (CIR) processes a query consisting of a reference image and a modification text, aiming to retrieve target images that not only resemble the reference image visually but also reflect the modification described in the caption. Unlike traditional image retrieval methods that rely on a single modality, CIR integrates visual and textual information, enabling more nuanced and constraint-based query representations. This unique capability has garnered growing interest from researchers. Despite its potential, the field lacks a systematic review that comprehensively examines its advancements and trends. This paper seeks to fill this gap by providing a detailed review of CIR research developments over the past five years. It categorizes existing methods into supervised approaches, which leverage triplet-labeled data for model training, and zero-shot approaches, which utilize unlabeled data to address CIR challenges. Additionally, the widely used benchmark datasets and evaluation indicators are comprehensively introduced. A comparative analysis of state-of-the-art methods across five datasets is also conducted, providing insights into their strengths and limitations. Ultimately, this paper also outlines potential research directions for the future.
科研通智能强力驱动
Strongly Powered by AbleSci AI