频域
计算机科学
目标检测
人工智能
特征(语言学)
特征提取
融合
模式识别(心理学)
变压器
计算机视觉
卷积(计算机科学)
领域(数学分析)
代表(政治)
传感器融合
特征学习
对象(语法)
图像融合
特征向量
时域
电子工程
空间频率
频道(广播)
融合规则
数据挖掘
模式
特征检测(计算机视觉)
特征模型
作者
Wencong Wu,Xiuwei Zhang,Hanlin Yin,Shuying Dai,Hongxi Zhang,Yanning Zhang
标识
DOI:10.48550/arxiv.2511.10046
摘要
Visible-infrared object detection has gained sufficient attention due to its detection performance in low light, fog, and rain conditions. However, visible and infrared modalities captured by different sensors exist the information imbalance problem in complex scenarios, which can cause inadequate cross-modal fusion, resulting in degraded detection performance. \textcolor{red}{Furthermore, most existing methods use transformers in the spatial domain to capture complementary features, ignoring the advantages of developing frequency domain transformers to mine complementary information.} To solve these weaknesses, we propose a frequency domain fusion transformer, called FreDFT, for visible-infrared object detection. The proposed approach employs a novel multimodal frequency domain attention (MFDA) to mine complementary information between modalities and a frequency domain feed-forward layer (FDFFL) via a mixed-scale frequency feature fusion strategy is designed to better enhance multimodal features. To eliminate the imbalance of multimodal information, a cross-modal global modeling module (CGMM) is constructed to perform pixel-wise inter-modal feature interaction in a spatial and channel manner. Moreover, a local feature enhancement module (LFEM) is developed to strengthen multimodal local feature representation and promote multimodal feature fusion by using various convolution layers and applying a channel shuffle. Extensive experimental results have verified that our proposed FreDFT achieves excellent performance on multiple public datasets compared with other state-of-the-art methods. The code of our FreDFT is linked at https://github.com/WenCongWu/FreDFT.
科研通智能强力驱动
Strongly Powered by AbleSci AI