计算机科学
稳健性(进化)
人工智能
特征提取
特征(语言学)
航空影像
目标检测
RGB颜色模型
计算机视觉
背景(考古学)
模式识别(心理学)
传感器融合
特征检测(计算机视觉)
图像融合
特征学习
一般化
上下文模型
光学(聚焦)
深度学习
人工神经网络
模态(人机交互)
融合
噪音(视频)
机器学习
多通道交互
作者
Bei Cheng,Bowen Xu,Wenjie Gan,QingWang Wang
标识
DOI:10.1109/jstars.2025.3648023
摘要
Multimodal object detection plays a crucial role in all-weather and multi-scene applications of aerial imagery. Existing studies mainly focus on multimodal fusion and inter-level feature interaction during feature extraction, that is, correcting or enhancing dual-branch weights through fused features or multimodal interactions, while neglecting the supplementation of missing modality features. This limitation can lead to noise propagation across layers and a reduction in interaction capability caused by feature absence. In this paper, we propose a Multimodal Collaborative Interactive Soft Fusion Network (MCISFNet) for RGB-infrared aerial image object detection. The proposed method introduces a Saliency-guided Multimodal Soft Fusion Mechanism (SMSFM), which explicitly directs attention to and enhances critical regions, dynamically adjusts feature weights, and integrates complementary information to mitigate the problem of missing data in dual-branch representations. To address the complexity of aerial scenarios, we further develop a Multi-scale Interactive Gating Module (MIGM) that explicitly incorporates multi-scale contextual information, enabling fine-grained refinement of primary modality features and enhancing the discriminability of fused representations. Moreover, we design a Cross-modal Global Context Collaborative Modeling (CGCCM) strategy, in which a cross-modal shared branch is constructed to jointly perform context extraction and feature fusion. This collaborative design not only improves the alignment of deep semantic features but also ensures that the learned RGB and IR features are more consistent and complementary, while reducing computational cost. Extensive experiments conducted on three multimodal aerial image detection datasets (DroneVehicle, VEDAI, and ODinMJ) demonstrate the robustness and generalization capability of the proposed MCISFNet framework.
科研通智能强力驱动
Strongly Powered by AbleSci AI