计算机科学
人工智能
答疑
光学(聚焦)
图像(数学)
视觉注意
感知器
自然语言处理
词(群论)
机器学习
自然语言
模式识别(心理学)
人工神经网络
认知
语言学
物理
生物
神经科学
哲学
光学
作者
Jinmeng Wu,Fulin Ge,Pengcheng Shu,Lei Ma,Yanbin Hao
标识
DOI:10.1109/aicit55386.2022.9930294
摘要
Visual Question and Answer (VQA) refers to a typical multimodal problem in the fields of computer vision and natural language processing, which aims to give an open-ended question about an image that can be answered accurately. The currently existing visual question answer models inevitably introduce redundant and inaccurate visual information when exploring the rich interaction between complex image targets and texts, and they also fail to focus effectively on the targets in the scene. To address this problem, the Question-Driven Multiple Attention Model (QDMA) is proposed. Firstly, Faster R-CNN and LSTM are used to extract visual features of images and textual features of questions. Then we design a question-driven attention network to obtain question regions of interest in images so that the model can accurately target relevant targets in complex scenes. To establish intensive interaction between the image region of interest and the question word, the co-attentive network consisting of self-attentive and guided-attentive units is introduced. Finally, the correct answer is obtained by inputting question features and image features into an answer prediction module consisting of two-layer Multi-Layer Perceptron. On the VQA2.0 dataset, the suggested method is empirically compared with other methods. The results reveal that the model outperforms other methods, demonstrating the usefulness of the framework.
科研通智能强力驱动
Strongly Powered by AbleSci AI