计算机科学
目标检测
人工智能
特征提取
特征(语言学)
棱锥(几何)
遥感
卷积(计算机科学)
卷积神经网络
计算机视觉
模式识别(心理学)
变压器
像素
上下文图像分类
遥感应用
特征学习
高光谱成像
对象(语法)
预处理器
语义特征
深度学习
视觉对象识别的认知神经科学
探测器
相似性(几何)
作者
Yang Feng,Shilong Jing,Yuchen Zhao,Hengyi Lv,Yisa Zhang,Ming Sun
标识
DOI:10.1109/lgrs.2025.3633625
摘要
Remote sensing object detection is a crucial task in ground analysis. Currently, object detectors based on convolution and transformer frameworks have shown significant performance. However, there are three pressing issues that still need to be addressed: 1) Detection of diminutive remote sensing targets exhibiting high inter-class similarity and imbalanced foreground-background distribution presents significant challenges; 2) Conventional CNN architectures demonstrate limited capability in capturing long-range dependencies, while Transformer frameworks incur substantial computational overhead; 3) Simply Feature Pyramid Network (FPN) fails to fusing the fine-grained characteristics of small targets. As the result, this work firstly introduces the Feature Split Attention Module (FSAM), which incorporates maxpooling and spatial attention mechanisms to decouple foreground and background features while preserving critical edge information. Moreover, we propose Global Convolutional Mamba Module (GCMM) that leverages MambaV2 architecture with SSD mechanisms for global feature extraction, thereby enhancing the long-range semantic modeling capabilities. Furthermore, Bidirectional Feature Pyramid Network (BiFPN) is adopted to strengthen multi-scale feature extraction for diminutive targets. Finally, these plug-and-play modules can be easily integrated into various object detection architectures. Experimental results demonstrate that our MVMamba model achieves 34.6% and 50.7% mAP@0.5:0.95 on the VisDrone and DIOR remote sensing datasets respectively, outperforming all other state-of-the-art approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI