FreDFT: Frequency Domain Fusion Transformer for Visible-Infrared Object Detection

频域 计算机科学 目标检测 人工智能 特征(语言学) 特征提取 融合 模式识别(心理学) 变压器 计算机视觉 卷积(计算机科学) 领域(数学分析) 代表(政治) 传感器融合 特征学习 对象(语法) 图像融合 特征向量 时域 电子工程 空间频率 频道(广播) 融合规则 数据挖掘 模式 特征检测(计算机视觉) 特征模型
作者
Wencong Wu,Xiuwei Zhang,Hanlin Yin,Shuying Dai,Hongxi Zhang,Yanning Zhang
标识
DOI:10.48550/arxiv.2511.10046
摘要

Visible-infrared object detection has gained sufficient attention due to its detection performance in low light, fog, and rain conditions. However, visible and infrared modalities captured by different sensors exist the information imbalance problem in complex scenarios, which can cause inadequate cross-modal fusion, resulting in degraded detection performance. \textcolor{red}{Furthermore, most existing methods use transformers in the spatial domain to capture complementary features, ignoring the advantages of developing frequency domain transformers to mine complementary information.} To solve these weaknesses, we propose a frequency domain fusion transformer, called FreDFT, for visible-infrared object detection. The proposed approach employs a novel multimodal frequency domain attention (MFDA) to mine complementary information between modalities and a frequency domain feed-forward layer (FDFFL) via a mixed-scale frequency feature fusion strategy is designed to better enhance multimodal features. To eliminate the imbalance of multimodal information, a cross-modal global modeling module (CGMM) is constructed to perform pixel-wise inter-modal feature interaction in a spatial and channel manner. Moreover, a local feature enhancement module (LFEM) is developed to strengthen multimodal local feature representation and promote multimodal feature fusion by using various convolution layers and applying a channel shuffle. Extensive experimental results have verified that our proposed FreDFT achieves excellent performance on multiple public datasets compared with other state-of-the-art methods. The code of our FreDFT is linked at https://github.com/WenCongWu/FreDFT.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
勤奋完成签到 ,获得积分10
4秒前
文艺的熠彤完成签到,获得积分10
6秒前
今天不熬夜完成签到 ,获得积分10
7秒前
DOC_XIONG应助科研通管家采纳,获得10
8秒前
8秒前
Double2and9完成签到 ,获得积分10
8秒前
苏紫梗桔完成签到,获得积分10
9秒前
做实验的猫完成签到,获得积分0
13秒前
阿尼完成签到 ,获得积分10
17秒前
阳光初之完成签到 ,获得积分10
17秒前
小月顺利毕业版完成签到,获得积分10
18秒前
勤恳的板凳完成签到 ,获得积分10
19秒前
2012csc完成签到 ,获得积分0
21秒前
21秒前
苗条的一一完成签到,获得积分10
22秒前
liuchang完成签到 ,获得积分10
22秒前
诚心闭月发布了新的文献求助10
24秒前
Nexus应助岁月如酒采纳,获得10
24秒前
幽默的迎天完成签到,获得积分10
26秒前
若琦2026完成签到 ,获得积分10
26秒前
8D完成签到,获得积分10
26秒前
99876完成签到 ,获得积分10
26秒前
俏皮友桃完成签到,获得积分10
30秒前
deswin完成签到,获得积分10
33秒前
Royzeng完成签到 ,获得积分10
34秒前
诚心闭月完成签到 ,获得积分10
34秒前
光电发布了新的文献求助10
34秒前
冷傲迎梅完成签到 ,获得积分10
39秒前
沐泽完成签到 ,获得积分10
40秒前
maxthon完成签到,获得积分10
40秒前
wave完成签到,获得积分10
41秒前
热情高跟鞋完成签到,获得积分10
42秒前
橘生淮南完成签到,获得积分10
42秒前
苗苗043完成签到,获得积分10
42秒前
Michael_li完成签到,获得积分10
45秒前
深情安青应助光电采纳,获得10
47秒前
听流沙完成签到 ,获得积分10
47秒前
大江大河完成签到 ,获得积分10
50秒前
四叶草完成签到 ,获得积分10
50秒前
fenger111完成签到,获得积分10
55秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Navigating Normative Orders. Interdisciplinary Perspectives 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
A Case Study on Hotels as Noncongregate Emergency Living Accommodations for Returning Citizens 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7754508
求助须知:如何正确求助?哪些是违规求助? 9301048
关于积分的说明 20260358
捐赠科研通 7336962
什么是DOI,文献DOI怎么找? 3310879
关于科研通互助平台的介绍 2462112
邀请新用户注册赠送积分活动 2324165