计算机科学
关系(数据库)
人工智能
图像(数学)
关系抽取
图形
情态动词
模式识别(心理学)
匹配(统计)
卷积神经网络
源代码
自然语言处理
理论计算机科学
数据挖掘
数学
操作系统
统计
化学
高分子化学
作者
Fei Zhao,Chunhui Li,Zhen Wu,Shangyu Xing,Xinyu Dai
标识
DOI:10.1145/3503161.3548228
摘要
Multimodal Named Entity Recognition (MNER) aims to locate and classify named entities mentioned in a (text, image) pair. However, dominant work independently models the internal matching relations in a pair of image and text, ignoring the external matching relations between different (text, image) pairs inside the dataset, though such relations are crucial for alleviating image noise in MNER task. In this paper, we primarily explore two kinds of external matching relations between different (text, image) pairs, i.e., inter-modal relations and intra-modal relations. On the basis, we propose a Relation-enhanced Graph Convolutional Network (R-GCN) for the MNER task. Specifically, we first construct an inter-modal relation graph and an intra-modal relation graph to gather the image information most relevant to the current text and image from the dataset, respectively. And then, multimodal interaction and fusion are leveraged to predict the NER label sequences. Extensive experimental results show that our model consistently outperforms state-of-the-art works on two public datasets. Our code and datasets are available at https://github.com/1429904852/R-GCN.
科研通智能强力驱动
Strongly Powered by AbleSci AI