可解释性
计算机科学
人工智能
卷积神经网络
目标检测
模式识别(心理学)
块(置换群论)
图像融合
神经编码
骨干网
对象(语法)
编码(社会科学)
计算机视觉
图像(数学)
分割
图像分割
传感器融合
融合
深度学习
相似性(几何)
编码(内存)
解码方法
模态(人机交互)
人工神经网络
语义学(计算机科学)
源代码
卷积码
图像处理
反向
特征提取
能见度
可视化
作者
Gargi Panda,Soumitra Kundu,Saumik Bhattacharya,Aurobinda Routray
标识
DOI:10.1109/tpami.2025.3643898
摘要
Multi-modal image fusion (MMIF) enhances the information content of the fused image by combining the unique as well as common features obtained from different modality sensor images, improving visualization, object detection, and many more tasks. In this work, we introduce an interpretable network for the MMIF task, named FNet, based on an $\ell _{0}$ℓ0-regularized multi-modal convolutional sparse coding (MCSC) model. Specifically, for solving the $\ell _{0}$ℓ0-regularized CSC problem, we design a learnable $\ell _{0}$ℓ0-regularized sparse coding (LZSC) block in a principled manner through deep unfolding. Given different modality source images, FNet first separates the unique and common features from them using the LZSC block and then these features are combined to generate the final fused image. Additionally, we propose an $\ell _{0}$ℓ0-regularized MCSC model for the inverse fusion process. Based on this model, we introduce an interpretable inverse fusion network named IFNet, which is utilized during FNet's training. Extensive experiments show that FNet achieves high-quality fusion results across eight different MMIF datasets. Furthermore, we show that FNet enhances downstream object detection and semantic segmentation in visible-thermal image pairs. We have also visualized the intermediate results of FNet, which demonstrates the good interpretability of our network.
科研通智能强力驱动
Strongly Powered by AbleSci AI