计算机科学
人工智能
语义学(计算机科学)
匹配(统计)
判决
比例(比率)
自然语言处理
桥接(联网)
模式
利用
短语
模式识别(心理学)
统计
数学
计算机网络
社会科学
物理
计算机安全
量子力学
社会学
程序设计语言
作者
Wenhui Li,Yan Wang,Yuting Su,Xuanya Li,An-An Liu,Yongdong Zhang
标识
DOI:10.1109/tmm.2021.3128744
摘要
Image and sentence matching is a critical task to bridge the visual and textual discrepancy due to the heterogeneous modalities. Great progress has been made by exploring the coarse-grained relationships between images and sentences or fine-grained relationships between regions and words. However, how to fully excavate and exploit corresponding relations between these two modalities is still challenging. In this work, we propose a novel Multi-scale Fine-grained Alignments Network (MFA), which can effectively explore multi-scale visual-textual correspondences to facilitate bridging cross-modal discrepancy. Specifically, word-scale matching module is firstly utilized to mine the basic but fundamental correspondences between a single word and independent region. Then, we propose a phrase-scale matching module to explore the relations between objects with the constraint of attribute and corresponding region, which can further reserve more associated information. To cope with the complex interactions among multiple phrases and images, we design the relation-scale matching module to capture high-order semantics between two modalities. Moreover, each matching module includes visual aggregation and textual aggregations, which can ensure the bi-directional coupling of multi-scale semantics. Extensive qualitative and quantitative experiments on two challenging datasets including Flickr30 K and MSCOCO, show that the proposed method achieves superior performance compared with the existing methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI