计算机科学
情态动词
人工智能
自然语言处理
命名实体识别
图形
情报检索
理论计算机科学
工程类
化学
高分子化学
系统工程
任务(项目管理)
作者
Guohui Ding,Wenjing Tang,Zhaoyi Yuan
标识
DOI:10.1109/smc54092.2024.10831708
摘要
Multimodal Named Entity Recognition (MNER) is a task that leverages multimodal information (such as text and images) to identify named entities within social media text. Traditional MNER methods primarily rely on simple interactions between text annotations and visual features, thus overlooking the specific correspondence between text and visual objects. Additionally, irrelevant visual noise may interfere with the final recognition results. Therefore, this paper proposes a Entity label-guided Graph Fusion Multi-modal Named Entity Recognition approach(ELGF), which utilizes pre-defined entity label information from input text as a bridge. Firstly, entity label detection tasks are employed to obtain entity label information. Then, the entity label information is utilized as a bridge between the two modalities to construct a multimodal interaction graph. This graph is inputted into a graph neural network, where attention and gate mechanisms are applied to interactively fuse multimodal information. Finally, a Conditional Random Field (CRF) decoding is used to predict the final MNER label sequence. Extensive experimental results demonstrate that compared to mainstream methods, the proposed model achieves competitive recognition accuracy on public datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI