计算机科学
特征(语言学)
人工智能
棱锥(几何)
模式识别(心理学)
失真(音乐)
增采样
点云
计算机视觉
特征提取
特征向量
骨料(复合)
频道(广播)
目标检测
对象(语法)
图像拼接
点(几何)
块(置换群论)
膨胀(度量空间)
体素
卷积(计算机科学)
空间分析
特征学习
电流(流体)
国家(计算机科学)
采样(信号处理)
空间分割
粒度
比例(比率)
骨干网
作者
X Y Liu,Ke Xu,Xinjie Wang,Z J Liu,Hanyun Wang,Y Guo
标识
DOI:10.1016/j.isprsjprs.2026.04.019
摘要
While recent 3D object detection methods have achieved impressive overall performance, they generally struggle to maintain high precision for small and distant targets. On the one hand, existing 3D backbones often lack sufficient multi-scale and channel-wise discrimination, causing rich features of large proximal objects to overwhelm the faint signals of small distant targets during downsampling. On the other hand, although current Mamba-based architectures achieve high performance with linear computational complexity, their reliance on unidirectional scanning ignores bidirectional spatial relationships, leading to geometric distortion and the neglect of critical features for rare voxels. To address these issues, drawing inspiration from the mechanism of Gated Attention for adaptive feature modulation, we propose GateMamba, a novel 3D backbone composed of stacked GateMamba blocks equipped with diverse feature gated mixers. To mitigate the loss of fine-grained spatial details during hierarchical downsampling, we design the GateMamba block, which incorporates a dense feature pyramid structure composed of several nested GateMamba layers and a scale feature gated mixer to adaptively weight and aggregate multi-scale features. To address the spatial distortion caused by the unidirectional scanning of standard state space models, we introduce a spatial-channel feature gated mixer within the GateMamba layer to bidirectionally aggregate spatial contexts and recalibrate channel responses. To prevent the feature vanishing of sparse instances during strided downsampling operations, we proposed the dilation voxel generation strategy, which proactively synthesizes features for foreground placeholders aligned with the sampling stride. Extensive experiments on the KITTI, Waymo, ONCE and NuScenes datasets demonstrate that GateMamba outperforms current state-of-the-art methods, such as achieving 73.8% and 67.5% mAP on the KITTI and ONCE datasets, respectively. More importantly, GateMamba specifically improves the detection precision of small and distant objects, such as outperforming the baseline by 2.5% and 2.4% in Level 1 and Level 2 mAP for the cyclist category on the Waymo validation dataset. Ablation studies further validate the contributions of diverse feature gated mixers in enhancing feature discriminabilities.
科研通智能强力驱动
Strongly Powered by AbleSci AI