人工智能
计算机科学
残余物
Boosting(机器学习)
计算机视觉
特征(语言学)
编码器
模式识别(心理学)
网格
融合
频道(广播)
图像融合
特征向量
特征提取
解耦(概率)
失真(音乐)
偏移量(计算机科学)
卷积神经网络
人工神经网络
数字图像
线性子空间
冗余(工程)
算法
图像(数学)
传感器融合
图像质量
可视化
滤波器(信号处理)
图像压缩
自编码
作者
Yuheng Gu,Shijie Li,Fengbo Wu,Dazhong Wu,Qing Cao,Xin Chen,Jiazhong Zhang,Yiwei Lou
标识
DOI:10.1109/cscloud66326.2025.00045
摘要
To tackle problems of structural redundancy, detail loss, and limited cross-modal feature fusion in infrared and visible image fusion, we introduce a two-stage fusion network based on state-space modeling with Mamba. Specially, the proposed network uses a Transformer-CNN decoupled encoder to extract global semantic and local texture features, while the Mamba module dynamically fuses base and detail information. With its selective state-space modeling mechanism, Mamba captures cross-modal dependencies efficiently under linear computational complexity, boosting feature expressiveness and reducing redundancy. The final fused output is refined using channel mixing, residual convolution, and channel compression modules. Training follows a two-phase scheme to separately optimize feature decoupling reconstruction and the quality of the fused image. Experiments on the MSRS, TNO, and RoadScene datasets demonstrate superior performance across 8 metrics, with notable advantages in VIF, SSIM, and MI. The proposed approach holds significant promise for Multi-modal applications, including remote sensing, smart grid systems, and digital twin modeling.
科研通智能强力驱动
Strongly Powered by AbleSci AI