因果推理
计算机科学
虚假关系
因果推理
因果模型
视觉推理
人工智能
一致性(知识库)
推论
因果一致性
因果关系(物理学)
答疑
因果关系
自然语言
过程(计算)
启发式
自然语言处理
认知科学
机器学习
因果结构
正规化(语言学)
认知心理学
常识推理
可视化
发电机(电路理论)
语义学(计算机科学)
钥匙(锁)
构造(python库)
破译
定性推理
水准点(测量)
作者
Wei Li,Fuyun Deng,Z. Y. Li
标识
DOI:10.1016/j.engappai.2026.113765
摘要
Explainable Visual Question Answering (EVQA) aims to not only predict accurate answers to visual questions but also generate human-friendly multimodal explanations that reveal the underlying reasoning process. Despite significant progress, existing EVQA methods suffer from two critical limitations: (1) they often rely on spurious cross-modal correlations (e.g., linguistic biases or visual shortcuts) rather than genuine causal relations, leading to unreliable reasoning; (2) the consistency between predicted answers and generated explanations is compromised due to the lack of explicit modeling of their causal dependencies. To address these issues, we propose a Cross-Modal Causal Reasoning (CMCR) framework that integrates causal inference with multimodal learning to disentangle causal effects from spurious correlations and enforce answer-explanation consistency. Specifically, CMCR incorporates three key innovations: (1) Causal Intervention, which employs backdoor adjustment to eliminate linguistic biases and frontdoor adjustment to mitigate visual shortcut biases; (2) a Neural-Symbolic Explanation Generator designed to translate symbolic reasoning processes into natural language explanations, thereby enhancing process explainability; and (3) Variational Causal Inference, which enforces causal consistency between answers and explanations. Experiments on benchmark datasets demonstrate that CMCR outperforms state-of-the-art methods, achieving a 1.19% higher accuracy, a 1.05% higher grounding for explanation quality, and a 0.42% higher answer-explanation consistency. • We formalize EVQA within a structural causal model to distinguish genuine causal paths. • We design a causal intervention module to refine visual and linguistic representations. • We formulate variational causal inference to enforce the causal consistency. • We introduce a neural-symbolic explanation generator to generate multimodal explanation. • Experiments demonstrate that CMCR outperforms state-of-the-art methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI