计算机科学
冗余(工程)
图像融合
特征提取
人工智能
块(置换群论)
计算机视觉
自然语言
融合
过程(计算)
光学(聚焦)
特征(语言学)
模式识别(心理学)
图像(数学)
操作系统
光学
物理
哲学
语言学
数学
几何学
作者
Jundong Zhang,Kangjian He,Dan Xu,Hongzhen Shi
标识
DOI:10.1109/tce.2025.3526792
摘要
The objective of infrared and visible image fusion is to produce a fused image that encompasses significant objects and intricate textures. However, existing methods frequently prioritize the extraction of complementary information, often overlooking the detrimental effects of redundant features. Moreover, due to the absence of authentic fused images, traditional mathematically defined loss functions face challenges in accurately modeling the characteristics of fused images. To address these challenges, this paper utilizes CLIP to design a natural language-guided, low-redundancy feature infrared and visible image fusion network. On one hand, we designed a Partial Feature Extraction(PFE) block and a Spatial-Channel Reconstruction Screening(SCRS) block to effectively reduce redundant features and enhance the focus on critical features. Additionally, we leveraged the CLIP model to bridge the gap between images and natural language, innovatively crafting a language-driven loss function to guide the fusion process through linguistic expressions. Extensive experiments conducted on multiple public datasets demonstrate that this method outperforms existing advanced techniques in both visual quality and quantitative assessment. Moreover, it achieves superior detection accuracy compared to current methods, reaching an advanced level of performance. The source code will be released at https://github.com/VCMHE/CNLFusion.
科研通智能强力驱动
Strongly Powered by AbleSci AI