多光谱图像
计算机科学
人工智能
RGB颜色模型
代表(政治)
计算机视觉
联营
目标检测
对象(语法)
特征(语言学)
融合
模式识别(心理学)
传感器融合
图像融合
判别式
频道(广播)
特征提取
能见度
GSM演进的增强数据速率
空间分析
钥匙(锁)
人工神经网络
融合规则
卷积神经网络
作者
Zhengju Jia,Ming Hui,Jinshu Huang,Zufeng Fu,Tingxiao Pan,Jinghong Yao
标识
DOI:10.1088/1361-6501/ae73a0
摘要
Abstract Multispectral object detection leverages the complementary characteristics of visible (RGB) and infrared (IR) modalities to improve target perception under low illumination, occlusion, and other complex environmental conditions. However, existing methods often suffer from imbalanced cross-modal fusion, leading to incomplete information integration and insufficient exploitation of fine-grained complementary cues. To address these limitations, we propose FMAF-YOLO, a lightweight dual-branch network for multispectral object detection. In the backbone, a C3k2-MLCA module preserves coarse spatial structure via local average pooling and introduces a global–local dual-branch attention mechanism to enhance the representation of small objects and edge regions. In the neck, a multi-scale cross-modal fusion module performs adaptive inter-layer interaction and weighted fusion of RGB and IR features to improve fusion balance and complementarity. In addition, a channel-aware dynamic sampling module jointly models spatial and channel information during upsampling, reducing feature blurring and preserving fine-grained details. Experiments on two public multispectral benchmarks demonstrate the effectiveness of the proposed method. On the M 3 FD dataset, FMAF-YOLO improves mAP@0.5 and mAP@0.5:0.95 by 4.1 and 3.7 percentage points, respectively, over the YOLO11n baseline. On the DroneVehicle dataset, evaluated under the official public train/validation/test protocol, the corresponding gains are 2.7 and 2.4 percentage points. These results demonstrate that FMAF-YOLO achieves effective fine-grained feature representation and adaptive cross-modal fusion while maintaining competitive computational efficiency.
科研通智能强力驱动
Strongly Powered by AbleSci AI