计算机视觉
人工智能
计算机科学
目标检测
对象(语法)
传感器融合
融合
模式识别(心理学)
特征(语言学)
噪音(视频)
图像融合
人脸检测
Viola–Jones对象检测框架
钥匙(锁)
摘要
Detecting targets under adverse visual conditions—including darkness, haze, and heavy occlusion—remains unreliable when the system depends on a single sensing source. Visible-spectrum imagery offers abundant color and texture cues, yet its reliability rapidly deteriorates under poor lighting. In contrast, thermal imagery remains stable across illumination changes but sacrifices detailed appearance information. To leverage the heterogeneous strengths of both sensing streams, we introduce a structured hierarchical fusion architecture termed the Hierarchical Cross-Modality Fusion Module (HCMFM), for robust RGB–infrared object detection. The proposed module operates at the representation level and can be inserted between backbone and head components without altering the original detection pipeline and enables multi-scale cross-modality interaction in a unified manner. Specifically, it consists of three cooperative components: a Cross-Modality Fusion Block (CMFB) that captures global cross-modality dependencies through transformer-based token fusion, a Cross Attention Interactive Block (CAIB) that explicitly mines complementary information between modalities via cross-attention mechanisms, and an Adaptive Gate Fusion (AGF) module that dynamically balances different fusion pathways according to scene characteristics. This hierarchical design effectively enhances discriminative representations while suppressing cross-modality interference under challenging conditions. Results obtained from standard benchmark datasets show consistent performance gains of the proposed framework relative to existing RGB–infrared detection approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI