情态动词
计算机科学
情报检索
材料科学
高分子化学
作者
Tianshi Wang,Fengling Li,Lei Zhu,Jingjing Li,Zheng Zhang,Heng Tao Shen
出处
期刊:Proceedings of the IEEE
[Institute of Electrical and Electronics Engineers]
日期:2024-11-01
卷期号:112 (11): 1716-1754
被引量:47
标识
DOI:10.1109/jproc.2024.3525147
摘要
With the exponential surge in diverse multimodal data, traditional unimodal retrieval methods struggle to meet the needs of users seeking access to data across various modalities. To address this, cross-modal retrieval has emerged, enabling interaction across modalities, facilitating semantic matching, and leveraging complementarity and consistency between heterogeneous data. Although prior literature has reviewed the field of cross-modal retrieval, it suffers from numerous deficiencies in terms of timeliness, taxonomy, and comprehensiveness. This article conducts a comprehensive review of cross-modal retrieval’s evolution, spanning from shallow statistical analysis techniques to vision-language pretraining (VLP) models. Commencing with a comprehensive taxonomy grounded in machine learning paradigms, mechanisms, and models, this article delves deeply into the principles and architectures underpinning existing cross-modal retrieval methods. Furthermore, it offers an overview of widely used benchmarks, metrics, and performances. Lastly, this article probes the prospects and challenges that confront contemporary cross-modal retrieval, while engaging in a discourse on potential directions for further progress in the field. To facilitate the ongoing research on cross-modal retrieval, we develop a user-friendly toolbox and an open-source repository at https://cross-modal-retrieval.github.io.
科研通智能强力驱动
Strongly Powered by AbleSci AI