激光雷达
计算机科学
情态动词
计算机视觉
目标检测
人工智能
融合
传感器融合
对象(语法)
遥感
模式识别(心理学)
地理
语言学
哲学
化学
高分子化学
标识
DOI:10.1109/jiot.2025.3584797
摘要
The cross-modal methods with LiDAR-image fusion have shown significant advantages in the field of 3D object detection. The existing multi-modal fusion strategies (such as feature-level or RoI-level) are difficult to balance global representation consistency and localization precision. Feature-level fusion methods fuse the different modal features in the middle layer of feature extraction, but easily cause feature space misalignment at the global scale, while RoI-level fusion methods integrate LiDAR and image features on local RoI grids, but highly depends on the quality of generated 3D proposals in the first stage. To overcome these limitations, we propose a multi-level fusion paradigm that integrate both feature-level and RoI-level fusion strategies in an end-to-end network, achieving fine-grained fusion of LiDAR and RGB data from global to local scales. Firstly, each frame of 3D point cloud and its 2D image are separately fed into LiDAR and image backbones to extract the corresponding 3D and 2D features. For each voxel feature encoded with 3D sparse convolution in LiDAR backbone, we calculate its corresponding 2D image feature based on camera-LiDAR transformation, and fuse them into dense multi-modal 3D voxel features for subsequent Region Proposal Network (RPN). In addition, we predict each foreground voxel score with a voxel point segmentation module and weight it into the global voxel-image fusion process to better guide the cross-modal fusion of two types of data. In RoI refinement stage, we further adopt a multi-modal RoI pooling to aggregate the multi-modal 3D voxel and 2D image features by deformable cross attention onto each RoI grid point. Related experiments on KITTI, NuScenes and Waymo Open Datasets prove the superior efficiency and accuracy of our proposed method.
科研通智能强力驱动
Strongly Powered by AbleSci AI