多光谱图像
计算机科学
人工智能
计算机视觉
融合
图像融合
传感器融合
目标检测
对象(语法)
多维系统
模式识别(心理学)
图像(数学)
数学
语言学
数学分析
哲学
作者
Fan Yang,Binbin Liang,Wei Li,Jianwei Zhang
标识
DOI:10.1109/tcsvt.2024.3454631
摘要
Multispectral object detection has attracted increasing attention recently due to its superior detection capacity under various illumination conditions. The key challenge lies in the effective aggregation of multi-spectral features to derive highly discriminative representations. To address this challenge, we propose a novel Multidimensional Fusion Network (MMFN) to explore multi-modal information from local, global, and channel perspectives. Specifically, at the local level, local features of different modalities and their inter-relationships are captured by a window-shifted fusion. As a complement to the local information, we designed a global interaction module that facilitates the fusion of holistic, high-level semantic information spanning the entire image. We distillate the channel dependencies and complementarities between different modalities through cross-channel learning and generate the final fused representation. Comprehensive experiments conducted on three publicly available datasets provide compelling evidence validating the superiority of the proposed methodology. The results exhibit notable performance gains over state-of-the-art multispectral object detectors. Our code will be released.
科研通智能强力驱动
Strongly Powered by AbleSci AI