稳健性(进化)
计算机科学
目标检测
无人机
人工智能
计算机视觉
特征提取
探测器
变更检测
图像传感器
传感器融合
提取器
对象(语法)
特征(语言学)
编码(集合论)
模式识别(心理学)
对象类检测
视觉对象识别的认知神经科学
源代码
图像分割
作者
Xin Wen,Heng Yin,Kai Li,Wanying Nie,Jianxun Zhao,Kechen Song
标识
DOI:10.1109/tgrs.2025.3645820
摘要
RGB-T object detection is increasingly applied in drone surveillance, autonomous driving, and intelligent transportation due to its robustness in complex conditions. However, most existing methods assume well-aligned image pairs and often overlook misalignment between modalities, which arises from drone motion, viewpoint shifts, and sensor inconsistencies, thereby hindering detection accuracy. To address this limitation, we present CMAI-Det, a framework designed for object detection in unaligned RGB-T images. The approach employs a dual-stream extractor to strengthen the representational capacity of visible and thermal features. A modality-cooperative alignment module, using the thermal stream as reference, integrates multi-scale deformable convolutions and attention to align visible features. An adaptive fusion scheme is further introduced to balance modality contributions according to their reliability under varying conditions, enhancing feature robustness. Finally, a perception-driven deformable detection head improves discriminability by reinforcing spatial selectivity and structural adaptability, enabling precise modeling of diverse object appearances. Extensive experiments on the DVTOD dataset show that CMAI-Det surpasses 12 state-of-the-art RGB-T detectors in challenging drone scenarios, achieving an mAP of 86.1. It also performs strongly at standard thresholds, with AP50 of 90.0 and AP75 of 82.3, underscoring its robustness and effectiveness in complex detection tasks. The code is available at https://github.com/yinhaixu2000-coder/CMAI-Detection.
科研通智能强力驱动
Strongly Powered by AbleSci AI