特征(语言学)
红外线的
融合
混合(物理)
人工智能
传感器融合
特征提取
计算机科学
计算机视觉
模式识别(心理学)
材料科学
光学
物理
哲学
语言学
量子力学
作者
Zhao Cai,Yong Ma,Jun Huang,Zhanchuan Cai,Ge Wang,Fan Fan
标识
DOI:10.1109/jsen.2025.3571945
摘要
Infrared and visible image fusion aims to overcome the limitations of each modality by combining their complementary advantages, thereby enhancing the performance of high-level vision tasks. Many deep learning-based methods generate fused images with rich details and overall consistency by capturing global dependencies among features. However, these methods often neglect cross-modal semantic relationships and fail to capture global dependencies and semantic information in image features, limiting their effectiveness in improving high-level visual task performance. To address these challenges, we propose WaveFusion, a novel progressive fusion network based on the multi-layer perceptron (MLP) architecture with global receptive fields. Our method first divides the input image into a predefined number of patches and represents them as wave functions, which are defined by their real values and phases. Subsequently, by employing a pre-trained semantic segmentation network and a semantic loss function, WaveFusion effectively integrates semantic information from the image patches into the wave function representation and progressively merges the phases from both modalities. As a result, our method captures semantic information from image patches across various scenes, enabling dynamic adjustments to their fusion weights. Extensive experiments demonstrate that our method surpasses state-of-the-art methods in both visual effects and quantitative metrics. Furthermore, these fused images enhance the performance of the semantic segmentation network, resulting in an approximate 3.5% increase in mean Intersection over Union (mIoU). The source code and pre-trained model will be released at: https://github.com/zc617/WaveFusion.
科研通智能强力驱动
Strongly Powered by AbleSci AI