计算机科学
特征(语言学)
融合
对象(语法)
频域
人工智能
领域(数学分析)
遥感
目标检测
计算机视觉
基于对象
传感器融合
特征模型
模式识别(心理学)
实时计算
地质学
数学
软件
数学分析
哲学
程序设计语言
语言学
作者
Xu Sun,Yinhui Yu,Qing Cheng
标识
DOI:10.1080/2150704x.2024.2305177
摘要
Fusing complementary information of visible and infrared radiation modalities can improve object detection performance for unmanned aerial vehicle (UAV) remote sensing images under insufficient illumination conditions. Although previous works have conducted some studies in this field, they have rarely considered the adaptive ability of multimodal feature fusion, which limits the performance improvement space for multimodal detectors. To this end, we propose an adaptive multimodal feature fusion method with a frequency domain gate based on DINO (detection transformer with improved denoising anchor boxes), called multimodal DINO. In our approach, a multimodal feature encoder with underlying feature sharing is designed, which efficiently extracts common and differential features through RGB-guided infrared radiation data transformation. Additionally, an adaptive frequency domain gate is introduced to dynamically learn the degree of dependence on frequency-filtered features of each modality when processing different samples. We evaluate the proposed method on the two multimodal detection remote sensing image datasets, VEDAI and DroneVehicle. Extensive experiments demonstrate that our approach achieves superior performance compared to basic detectors and existing multimodal detection methods. Our code is available at https://github.com/cq100/multimodalDINO.
科研通智能强力驱动
Strongly Powered by AbleSci AI