遥感
桥接(联网)
计算机科学
目标检测
人工智能
计算机视觉
地球遥感
遥感应用
模式(计算机接口)
高光谱成像
对象(语法)
卫星广播
合成孔径雷达
激光雷达
深度学习
基于对象
图像处理
作者
Lulu Meng,Xingxing Ju,Zhi Li
标识
DOI:10.1109/tgrs.2026.3697288
摘要
Visible-infrared (RGB-T) multimodal object detection on unmanned aerial vehicles is challenged by physical parallax, modality conflicts, and information loss of small targets (sub-32 pixels). Existing fusion methods struggle with precise alignment in dynamic environments and suffer from representational competition, limiting detection accuracy. To address these issues, this paper proposes a parallax-aware detection framework, PAO-YOLO, built on a deep semantic bridging architecture. It introduces a Modality-Invariant Feature Disentanglement (MIFD) module to decouple geometric structures from modality-specific noise within an orthogonal common subspace. These structural priors then guide an Elastic Deformation Adaptive Fusion (EDAF) module, which employs predicted offset fields and deformable convolutions for pixel-level alignment. Additionally, a high-resolution P2 detection layer preserves spatial gradients for enhanced small-target perception. Experiments on the DroneVehicle and DVTOD datasets show that PAO-YOLO achieves 82.2% and 89.2% mAP@0.5, respectively.The method effectively mitigates parallax artifacts and improves sensitivity to small targets, sustaining a robust inference speed of 85.6 Frames Per Second (FPS) on a standard GPU. Code is available at: https://github.com/PAO-YOLO-Team/PAO-YOLO.git.
科研通智能强力驱动
Strongly Powered by AbleSci AI