计算机科学
情态动词
人工智能
图像检索
图像(数学)
计算机视觉
情报检索
化学
高分子化学
作者
Zhangxiang Shi,Yunlai Ding,Junyu Dong,Tianzhu Zhang
标识
DOI:10.1109/tcsvt.2025.3605958
摘要
Existing image-text retrieval methods mainly rely on region and word features to measure cross-modal similarities. Thus, dense cross-modal semantic alignment which matches regions and words becomes crucial. However, this is non-trivial due to the heterogeneity gap and the cross-modal attention used to achieve this alignment is inefficient. Towards solving this problem, we propose a novel framework that goes beyond the previous one-tower and two-tower frameworks to learn cross-modal consensus efficiently. The proposed framework does not align regions and words directly like existing methods but uses semantic prototypes as a bridge to attend specific contents with the same semantics among different modalities through semantic decoders, through which cross-modal semantic alignment is naturally achieved. Furthermore, we design a novel plug-and-play self-correction method based on optimal transport to alleviate the drawbacks of incomplete pairwise labels in existing multimodal datasets. On top of various base backbones, we carry out extensive experiments on two benchmark datasets, i.e., Flickr30K and MS-COCO, demonstrating the effectiveness, superiority and generalization of our method.
科研通智能强力驱动
Strongly Powered by AbleSci AI