Visual Chain-of-Thought Prompting for Knowledge-Based Visual Reasoning

视觉思维视觉推理认知科学心理学认知心理学计算机科学神经科学

作者

Zhenfang Chen,Qinhong Zhou,Yikang Shen,Yining Hong,Zhiqing Sun,Dan Gutfreund,Chuang Gan

出处

期刊：Proceedings of the ... AAAI Conference on Artificial Intelligence [Association for the Advancement of Artificial Intelligence (AAAI)]
日期：2024-03-24 卷期号：38 (2): 1254-1262 被引量：20

链接

aaai.orgdoi.org

标识

DOI：10.1609/aaai.v38i2.27888

摘要

Knowledge-based visual reasoning remains a daunting task since it not only requires machines to interpret the concepts and relationships from visual scenes but also associate them with external world knowledge to conduct a chain of reasoning on open-world questions. Previous works, however, treat visual perception and language-based reasoning as two independent modules, failing to attend to both modules throughout all stages of reasoning. To this end, we propose Visual Chain-of-thought Prompting (VCTP) for knowledge-based reasoning, which involves the interaction between visual content and natural language in an iterative step-by-step reasoning manner. VCTP contains three stages, see, think, and confirm. The see stage scans the image and grounds the visual concept candidates with a visual perception model. The think stage adopts a pre-trained large language model (LLM) to attend to key visual concepts from natural language questions adaptively. It then transforms key visual context into text context for prompting with a visual captioning model, and adopts the LLM to generate the answer. The confirm stage further uses the LLM to generate the supporting rationale to the answer, which is then passed through a cross-modality classifier to verify that it’s consistent with the visual context. We iterate through the think-confirm stages to ensure the verified rationale is consistent with the answer. We conduct experiments on a range of knowledge-based visual reasoning datasets. We found our VCTP enjoys several benefits, 1). it achieves better performance than the previous few-shot learning baselines; 2). it enjoys the total transparency and trustworthiness of the whole reasoning process by providing rationales for each reasoning step; 3). it is computation-efficient compared with other fine-tuning baselines. Our code is available at https://github.com/UMass-Foundation-Model/VisualCoT.git

求助该文献

最长约 10秒，即可获得该文献文件

Visual Chain-of-Thought Prompting for Knowledge-Based Visual Reasoning

今日热心研友