模态(人机交互)
人工智能
计算机科学
融合
先验概率
图像(数学)
计算机视觉
图像融合
贝叶斯概率
语言学
哲学
作者
Guanyao Wu,Haoyu Liu,Hongming Fu,Yichuan Peng,Jinyuan Liu,Xin Fan,Risheng Liu
出处
期刊:
日期:2025-06-10
卷期号:: 17882-17891
被引量:18
标识
DOI:10.1109/cvpr52734.2025.01666
摘要
Multi-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting to downstream tasks remains challenging. Recent approaches attempt task-specific design but rarely achieve "The Best of Both Worlds" due to inconsistent optimization goals. To address these issues, we propose a novel method that leverages the semantic knowledge from the Segment Anything Model (SAM) to Grow the quality of fusion results and Enable downstream task adaptability, namely SAGE. Specifically, we design a Semantic Persistent Attention (SPA) Module that efficiently maintains source information via the persistent repository while extracting high-level semantic priors from SAM. More importantly, to eliminate the impractical dependence on SAM during inference, we introduce a bi-level optimization-driven distillation mechanism with triplet losses, which allow the student network to effectively extract knowledge. Extensive experiments show that our method achieves a balance between high-quality visual results and downstream task adaptability while maintaining practical deployment efficiency. The code is available at https://github.com/RollingPlain/SAGE_IVIF.
科研通智能强力驱动
Strongly Powered by AbleSci AI