计算机科学
深度学习
人工智能
人工神经网络
计算机视觉
作者
Ma Qin,Lin Zeng,Shikui Tu,Lei Xu
出处
期刊:
日期:2024-12-03
卷期号:: 3592-3597
标识
DOI:10.1109/bibm62325.2024.10822726
摘要
Causal intervention has been widely used in deep learning to tackle the confounding problems when the data are out of distribution. A representative class of causal-intervention-based strategy is invariant risk minimization, yet manually annotating is indispensable to indicate the environmental splits. A feasible solution is to use a pair of complementary attention, one of which focuses on the foreground target and another extracts the background features. The environments are then split unsupervisedly. However, in terms of medical images, the background is not as simple as normal images (e.g., a camel in desert), and using all the background feature as confounder is redundant and will degrade the expected generalization performance. In this paper, we develop a novel multimodal deep neural network which automatically extracts the confounded background features of chest X-ray (CXR) images by a multimodal approach, i.e., feature fusion with explicit confounders such as demographic information. To achieve this, we design an architecture with two modules, a classifier module taking images as input and an environment split module taking demographic tables as input. The cross attention is then conducted between the encoded background features and demographic features to extract the actual confounded background features. The proposed method meaningfully improves the generalization performance on multiple CXR datasets and accurately locates the lesions. The source code is available at https://github.com/CMACH508/CausalCXR.
科研通智能强力驱动
Strongly Powered by AbleSci AI