模块化设计
计算机科学
答疑
水准点(测量)
注意力网络
钥匙(锁)
人工智能
深度学习
集合(抽象数据类型)
视觉注意
机器学习
自然语言处理
情报检索
心理学
程序设计语言
感知
计算机安全
大地测量学
操作系统
神经科学
地理
作者
Yu Zhou,Jun Yu,Yuhao Cui,Dacheng Tao,Qi Tian
标识
DOI:10.1109/cvpr.2019.00644
摘要
Visual Question Answering (VQA) requires a fine-grained and simultaneous understanding of both the visual content of images and the textual content of questions. Therefore, designing an effective `co-attention' model to associate key words in questions with key objects in images is central to VQA performance. So far, most successful attempts at co-attention learning have been achieved by using shallow models, and deep co-attention models show little improvement over their shallow counterparts. In this paper, we propose a deep Modular Co-Attention Network (MCAN) that consists of Modular Co-Attention (MCA) layers cascaded in depth. Each MCA layer models the self-attention of questions and images, as well as the question-guided-attention of images jointly using a modular composition of two basic attention units. We quantitatively and qualitatively evaluate MCAN on the benchmark VQA-v2 dataset and conduct extensive ablation studies to explore the reasons behind MCAN's effectiveness. Experimental results demonstrate that MCAN significantly outperforms the previous state-of-the-art. Our best single model delivers 70.63% overall accuracy on the test-dev set.
科研通智能强力驱动
Strongly Powered by AbleSci AI