计算机科学
推论
答疑
人工智能
领域知识
知识图
情态动词
图形
自然语言处理
领域(数学分析)
场景图
情报检索
理论计算机科学
数学分析
化学
高分子化学
渲染(计算机图形)
数学
作者
Longbao Wang,Jinhao Zhang,Libing Zhang,S. H. Zhang,Shufang Xu,Lin Yu,Hongmin Gao
标识
DOI:10.1142/s1469026824500342
摘要
Knowledge-based visual question answering relies on open-ended external knowledge and a fine-grained comprehension of both the visual content of images and semantic information. Existing methods for utilizing knowledge have the following limitations: (1) Language pre-training methods output answers in the form of plain text, which only understand shallow visual content; (2) The knowledge retrieved by image objects as labels is represented as first-order logic, making it difficult to infer complex questions. To address the above problems, this paper integrates visual-textual multimodal information, accumulates domain-specific and external multi-modal knowledge, introduces and supplements external objective facts, and proposes a multimodal knowledge graph construction and fact-assisted reasoning network (MKGFA). The network consists of three parts: the multimodal knowledge graph construction module (MKGC), the objective fact-assisted reasoning module (FAR), and the answer inference module. The MKGC engages in the coarse-to-fine-grained learning of triplet representations for multimodal knowledge units. The FAR establishes deep cross-modal relations between visual objects and factual words for correlating real answers. The answer inference module makes the final decision based on the results of both. Among them, the former two modules employ a pre-training and fine-tuning strategy, systematically accumulating foundational and domain-specific knowledge. Compared with the state-of-the-arts, MKGFA achieves 1.09% and 0.7% higher accuracy on the two challenging OKVQA and KRVQA datasets, respectively. The experimental results demonstrate the complementary advantages of the integration of the two modules.
科研通智能强力驱动
Strongly Powered by AbleSci AI