融合
频域
特征(语言学)
目标检测
补偿(心理学)
计算机科学
人工智能
RGB颜色模型
对象(语法)
传感器融合
领域(数学分析)
计算机视觉
模式识别(心理学)
数学
心理学
数学分析
哲学
语言学
精神分析
作者
Yinbo Gao,Zhuhua Liao,Yizhi Liu,Aiping Yi,Guo‐Qiang Zhang
标识
DOI:10.1109/jsen.2025.3559057
摘要
The multi-modal object detection technology based on visible–thermal vision sensors has drawn significant attention as it is capable of achieving reliable object detection in complex scenes with challenging lighting conditions such as low light or backlight. However, there has been a lack of focus on the frequency domain feature information of the visible and thermal modalities themselves, as well as the complementarity of cross-modal features. Furthermore, current visible and thermal feature fusion methods only utilize feature information from the current layer, neglecting context information. Therefore, this paper proposes a novel network framework for enhancing frequency domain Characteristics and fusing cross-layer and cross-modal features to compensate for these limitations. This framework introduces two key modules: the Frequency Domain Characteristics Enhancement (FCE) module and the Cross-Layer and Cross-Modal Feature Compensation Fusion (CCF) module. The FCE module consists of two sub-modules. Reduce high-frequency information loss(FCE-RHFL) module operates on the visible modality to reduce high-frequency information loss, utilizing methods such as high-frequency masking and frequency recombination. Meanwhile, Enhance high-frequency information representation(FCE-EHFR) module enhances high-frequency information representation for the thermal modality through convolutions with different kernels and frequency domain enhancement techniques. In the CCF module, cross-modal and cross-layer feature compensation methods are employed to compensate for differences in modalities, followed by capture complementary information across modalities using a query-guided cross-attention mechanism. Finally, we conduct experimental comparisons on the KAIST and FLIR datasets, and the experimental results demonstrate that our method has excellent performance and robust detection results.
科研通智能强力驱动
Strongly Powered by AbleSci AI