计算机视觉
人工智能
计算机科学
RGB颜色模型
目标检测
特征(语言学)
对象(语法)
传感器融合
探测器
特征提取
芯(光纤)
图像分辨率
可视化
图像融合
变更检测
模式识别(心理学)
融合
卫星
模式
高分辨率
作者
Cao, Shuyu,Chen, Minxin,Song, Yucheng,Chen, Zhaozhong,Zhang, Xinyou
标识
DOI:10.48550/arxiv.2511.19134
摘要
Small object detection in Unmanned Aerial Vehicle (UAV) imagery is a persistent challenge, hindered by low resolution and background clutter. While fusing RGB and infrared (IR) data offers a promising solution, existing methods often struggle with the trade-off between effective cross-modal interaction and computational efficiency. In this letter, we introduce MambaRefine-YOLO. Its core contributions are a Dual-Gated Complementary Mamba fusion module (DGC-MFM) that adaptively balances RGB and IR modalities through illumination-aware and difference-aware gating mechanisms, and a Hierarchical Feature Aggregation Neck (HFAN) that uses a ``refine-then-fuse'' strategy to enhance multi-scale features. Our comprehensive experiments validate this dual-pronged approach. On the dual-modality DroneVehicle dataset, the full model achieves a state-of-the-art mAP of 83.2%, an improvement of 7.9% over the baseline. On the single-modality VisDrone dataset, a variant using only the HFAN also shows significant gains, demonstrating its general applicability. Our work presents a superior balance between accuracy and speed, making it highly suitable for real-world UAV applications.
科研通智能强力驱动
Strongly Powered by AbleSci AI