计算机科学
超图
情态动词
卷积(计算机科学)
动量(技术分析)
理论计算机科学
情报检索
人工智能
数学
离散数学
人工神经网络
财务
经济
化学
高分子化学
摘要
Cross-modal retrieval tasks, encompassing the retrieval of image–text, video–audio, and more, are progressively gaining significance in response to the exponential growth of information on the Internet. However, there has always been a cloud hanging over multimodal tasks due to the inherent challenges in aligning different modalities with distinct physical meanings. Most previous works simply rely on a single multimodal encoder or a novel similarity calculation for fusion, which often result in unsatisfactory performance. To tackle this challenge, we introduce a Momentum Hypergraph Convolutional Network (MoHGCN) for multimodal representation learning, which strengthens the alignment of both visual and textual data before the retrieval process. Specifically, MoHGCN utilizes contrastive learning to select the most challenging negative and positive samples to form hyperedges and completes the modality alignment through two rounds of fusion. Subsequently, the fully integrated node features and global features are fused using a fusion encoder to obtain the final multimodal representation vector for image–text retrieval. Extensive experiments are conducted on two widely used datasets, namely Flickr30K and MSCOCO, to demonstrate the superiority of the proposed MoHGCN approach in achieving the state-of-the-art performances.
科研通智能强力驱动
Strongly Powered by AbleSci AI